Earlier  
Posted Nick Remark
#openstack-nova - 2022-08-31
14:13:02 dansmith gmann: yeah I mean whatever check we're doing that gives that error
14:13:15 dansmith is there a tempest test that is disabling a compute by chance for testing?
14:13:18 sean-k-mooney its proably worth checkign the schduler logs
14:13:33 dansmith oh is this just the scheduler "is this host okay" check?
14:15:20 gmann I do not think we have any such test of touching the compute services enable/disable
14:15:30 gibi it is not about sevice state
14:15:39 gibi the unshelve fails on compute_node_get_all_by_host
14:15:48 gibi so the compute is not in the DBV
14:15:49 gibi DB
14:17:15 gibi https://github.com/openstack/nova/blob/master/nova/compute/api.py#L4577
14:17:21 gibi this where the unshelve fails
14:18:53 sean-k-mooney so it is acutlly trying to shelve and unshlve multiple times
14:19:15 sean-k-mooney it shleve offloaded and shelve to the compute
14:19:18 sean-k-mooney then it shelved again
14:19:28 sean-k-mooney and failed to unshleve to the contoler with the not found issue
14:19:52 gibi yes
14:20:12 gibi first it unselves back to where it was, then it unshelves to the other compute
14:20:42 sean-k-mooney and req-b3e7705d-a20b-42b3-8eb5-594181a096c8 is the request that fialed
14:20:52 sean-k-mooney which failed in the api right
14:20:57 sean-k-mooney its not in the secheuler
14:21:02 gibi yes
14:21:07 gmann yes it is failing in 2nd time unshelve on another host
14:21:10 gibi it is not fails on the scheduler
14:21:16 gibi it fails in the compute api
14:21:17 gibi https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/controller/logs/screen-n-api.txt
14:21:20 gibi https://github.com/openstack/nova/blob/master/nova/compute/api.py#L4577
14:21:33 gibi unshelve checks is the requested host exists
14:21:49 gibi before it goes to the conductor
14:21:56 sean-k-mooney by calling compute_node_get_all_by_host
14:22:20 gibi it calls get_first_node_by_host_for_old_compat
14:22:26 sean-k-mooney thats on the main db i.e. cell db right
14:22:39 dansmith so that is looking up by the service, and getting the first compute node on the service
14:22:45 dansmith not by_host_and_node
14:23:13 sean-k-mooney for libvirt there should only be one but ya
14:23:32 sean-k-mooney if shelve was a thing for vmware or ironic that would not be ideal
14:23:38 dansmith sean-k-mooney: right but this is very loose association which doesn't have FK constraints
14:24:05 sean-k-mooney why is it hitting the cell db. it could look at the host mapping table in the api db right
14:24:14 sean-k-mooney or cell mappings table
14:24:21 dansmith eh?
14:24:29 dansmith those just tell us which db the host is in
14:24:43 sean-k-mooney right but if all you care about is does it exist is that not enough
14:24:51 gibi sean-k-mooney: you have a point, so the compute-api does a cell DB query without targeting the cell first
14:24:51 dansmith nothing about whether it's really there or up or whatever
14:25:23 sean-k-mooney dansmith: i guess it could be down but im not sure this is checkign that
14:25:28 dansmith sean-k-mooney: or deleted
14:25:48 sean-k-mooney i would expct that to be check by the schduler to be honest
14:25:53 dansmith gibi: these hosts are in the same cell though so if that wasn't working, we wouldn't have found the first one right?
14:26:22 gibi this is the multicell job so I assume we have hosts in different cells
14:26:45 sean-k-mooney yes there are only two nodes in the job
14:26:51 sean-k-mooney so each is in a different cell
14:26:59 dansmith multicell meaning superconductor not actually one host per cell right?
14:27:27 sean-k-mooney well we can check but i tought this was 2 cells
14:27:28 dansmith or is this really unshelving to different cells?
14:27:50 dansmith can we unshelve across cells?
14:27:52 dansmith I thought not
14:28:16 dansmith so maybe we're targeted at the cell where the instance is and not the cell where the host is that we're unshelving to?
14:28:24 dansmith not sure how that would have ever passed
14:28:48 sean-k-mooney the compute is not runnign a second conductor
14:28:55 gibi I'm looking at the compute.api.API.unshelve and I think the context is not targeted
14:29:14 sean-k-mooney oh the contoler is ruuning both cell1 and cell2 conductors
14:29:36 sean-k-mooney and the supper conductor
14:30:24 sean-k-mooney https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/compute1/logs/etc/nova/nova_cell1_conf.txt
14:30:51 sean-k-mooney transport_url = rabbit://stackrabbit:secretrabbit@10.208.192.130:5672/nova_cell2
14:31:00 sean-k-mooney so ya compute1 is nova_cell2
14:31:20 sean-k-mooney and contoler is transport_url = rabbit://stackrabbit:secretrabbit@10.208.192.130:5672/nova_cell1
14:31:26 gibi I don't know why this ever passed the test
14:31:45 dansmith I also don't know how unshelve could be using an untargeted context
14:31:54 dansmith it should be pointing at cell0 all the time which would never work
14:32:32 sean-k-mooney dansmith: did you say this just started failing recently?
14:32:42 sean-k-mooney because we merged unshelve to host a while ago
14:32:43 dansmith sean-k-mooney: it just recently merged I think gmann said
14:32:50 sean-k-mooney the tempest test?
14:33:16 sean-k-mooney if so i guess we dont run the multi-cell job on tempest
14:33:44 gmann dansmith: sean-k-mooney gibi that passed in tempest-multinode-full-py3 job and I do not think we checked multi-cell job run for this
14:33:56 dansmith the base nova conf doesn't have a db connection pointing at cell0, which I assume is what the apis are using
14:34:02 gmann yeah we do not run multi-cell job there and only checked tempest-multinode-full-py3 job passing it
14:34:22 sean-k-mooney gmann: ack that makes sense
14:34:32 gibi yeah nova-multicell did not run on the tempest patch
14:34:41 sean-k-mooney ok so we have two thing one we should skip this temporaly on the multi-cell-job
14:34:44 sean-k-mooney and then fix the bug
14:34:56 gibi yes
14:34:58 dansmith a normal run should have the api pointing at cell0: https://storage.bhs.cloud.ovh.net/v1/AUTH_dcaab5e32b234d56b626f72581e3644c/zuul_opendev_logs_8cf/831219/22/check/tempest-integrated-compute/8cfb267/controller/logs/etc/nova/nova_conf.txt
14:35:03 dansmith but that multicell run has nothing defined there
14:35:10 sean-k-mooney we at least shoudl check we are not unshelving across cells for now but that actully should work
14:35:18 sean-k-mooney since we use shelve for
14:35:21 sean-k-mooney cross cell resize
14:35:28 dansmith sean-k-mooney: it shouldn't work
14:35:38 dansmith because cross-cell migration does other stuff on top of the shelve
14:35:53 sean-k-mooney oh your right
14:36:22 gmann and cross-cell is not in that feature acope right? which is added in this cycle only.
14:36:24 gmann scope
14:36:33 gmann sean-k-mooney: agree to skip it in multi-cell job
14:36:38 sean-k-mooney it was not explcitly no
14:36:56 sean-k-mooney so we can just document the limitation i guess
14:37:05 gmann yes
14:37:05 gmann but will be good to add in documnt/releasenotes
14:38:07 sean-k-mooney dansmith: ohter then usign multiple port binding to test if the destination cell can bind the port. how much more does cross cell migration do over shelve. there is a bunch of stuff in the supper conductor to copy the instnace object between the cell db right
14:38:33 dansmith yeah it moves the thing between DBs, that's the big thing
14:39:13 sean-k-mooney well for now i guess its fine the error could be improved but the api wont let you currpt the db or anything like that
14:39:27 dansmith so this multicell job has two cells and two computes, one compute per cell? it must be disabling live migration and lots of other stuff that would normally require multinode right/

Earlier   Later