| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-08-31 | |||
| 14:17:21 | gibi | this where the unshelve fails | |
| 14:18:53 | sean-k-mooney | so it is acutlly trying to shelve and unshlve multiple times | |
| 14:19:15 | sean-k-mooney | it shleve offloaded and shelve to the compute | |
| 14:19:18 | sean-k-mooney | then it shelved again | |
| 14:19:28 | sean-k-mooney | and failed to unshleve to the contoler with the not found issue | |
| 14:19:52 | gibi | yes | |
| 14:20:12 | gibi | first it unselves back to where it was, then it unshelves to the other compute | |
| 14:20:42 | sean-k-mooney | and req-b3e7705d-a20b-42b3-8eb5-594181a096c8 is the request that fialed | |
| 14:20:52 | sean-k-mooney | which failed in the api right | |
| 14:20:57 | sean-k-mooney | its not in the secheuler | |
| 14:21:02 | gibi | yes | |
| 14:21:07 | gmann | yes it is failing in 2nd time unshelve on another host | |
| 14:21:10 | gibi | it is not fails on the scheduler | |
| 14:21:16 | gibi | it fails in the compute api | |
| 14:21:17 | gibi | https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/controller/logs/screen-n-api.txt | |
| 14:21:20 | gibi | https://github.com/openstack/nova/blob/master/nova/compute/api.py#L4577 | |
| 14:21:33 | gibi | unshelve checks is the requested host exists | |
| 14:21:49 | gibi | before it goes to the conductor | |
| 14:21:56 | sean-k-mooney | by calling compute_node_get_all_by_host | |
| 14:22:20 | gibi | it calls get_first_node_by_host_for_old_compat | |
| 14:22:26 | sean-k-mooney | thats on the main db i.e. cell db right | |
| 14:22:39 | dansmith | so that is looking up by the service, and getting the first compute node on the service | |
| 14:22:45 | dansmith | not by_host_and_node | |
| 14:23:13 | sean-k-mooney | for libvirt there should only be one but ya | |
| 14:23:32 | sean-k-mooney | if shelve was a thing for vmware or ironic that would not be ideal | |
| 14:23:38 | dansmith | sean-k-mooney: right but this is very loose association which doesn't have FK constraints | |
| 14:24:05 | sean-k-mooney | why is it hitting the cell db. it could look at the host mapping table in the api db right | |
| 14:24:14 | sean-k-mooney | or cell mappings table | |
| 14:24:21 | dansmith | eh? | |
| 14:24:29 | dansmith | those just tell us which db the host is in | |
| 14:24:43 | sean-k-mooney | right but if all you care about is does it exist is that not enough | |
| 14:24:51 | dansmith | nothing about whether it's really there or up or whatever | |
| 14:24:51 | gibi | sean-k-mooney: you have a point, so the compute-api does a cell DB query without targeting the cell first | |
| 14:25:23 | sean-k-mooney | dansmith: i guess it could be down but im not sure this is checkign that | |
| 14:25:28 | dansmith | sean-k-mooney: or deleted | |
| 14:25:48 | sean-k-mooney | i would expct that to be check by the schduler to be honest | |
| 14:25:53 | dansmith | gibi: these hosts are in the same cell though so if that wasn't working, we wouldn't have found the first one right? | |
| 14:26:22 | gibi | this is the multicell job so I assume we have hosts in different cells | |
| 14:26:45 | sean-k-mooney | yes there are only two nodes in the job | |
| 14:26:51 | sean-k-mooney | so each is in a different cell | |
| 14:26:59 | dansmith | multicell meaning superconductor not actually one host per cell right? | |
| 14:27:27 | sean-k-mooney | well we can check but i tought this was 2 cells | |
| 14:27:28 | dansmith | or is this really unshelving to different cells? | |
| 14:27:50 | dansmith | can we unshelve across cells? | |
| 14:27:52 | dansmith | I thought not | |
| 14:28:16 | dansmith | so maybe we're targeted at the cell where the instance is and not the cell where the host is that we're unshelving to? | |
| 14:28:24 | dansmith | not sure how that would have ever passed | |
| 14:28:48 | sean-k-mooney | the compute is not runnign a second conductor | |
| 14:28:55 | gibi | I'm looking at the compute.api.API.unshelve and I think the context is not targeted | |
| 14:29:14 | sean-k-mooney | oh the contoler is ruuning both cell1 and cell2 conductors | |
| 14:29:36 | sean-k-mooney | and the supper conductor | |
| 14:30:24 | sean-k-mooney | https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/compute1/logs/etc/nova/nova_cell1_conf.txt | |
| 14:30:51 | sean-k-mooney | transport_url = rabbit://stackrabbit:secretrabbit@10.208.192.130:5672/nova_cell2 | |
| 14:31:00 | sean-k-mooney | so ya compute1 is nova_cell2 | |
| 14:31:20 | sean-k-mooney | and contoler is transport_url = rabbit://stackrabbit:secretrabbit@10.208.192.130:5672/nova_cell1 | |
| 14:31:26 | gibi | I don't know why this ever passed the test | |
| 14:31:45 | dansmith | I also don't know how unshelve could be using an untargeted context | |
| 14:31:54 | dansmith | it should be pointing at cell0 all the time which would never work | |
| 14:32:32 | sean-k-mooney | dansmith: did you say this just started failing recently? | |
| 14:32:42 | sean-k-mooney | because we merged unshelve to host a while ago | |
| 14:32:43 | dansmith | sean-k-mooney: it just recently merged I think gmann said | |
| 14:32:50 | sean-k-mooney | the tempest test? | |
| 14:33:16 | sean-k-mooney | if so i guess we dont run the multi-cell job on tempest | |
| 14:33:44 | gmann | dansmith: sean-k-mooney gibi that passed in tempest-multinode-full-py3 job and I do not think we checked multi-cell job run for this | |
| 14:33:56 | dansmith | the base nova conf doesn't have a db connection pointing at cell0, which I assume is what the apis are using | |
| 14:34:02 | gmann | yeah we do not run multi-cell job there and only checked tempest-multinode-full-py3 job passing it | |
| 14:34:22 | sean-k-mooney | gmann: ack that makes sense | |
| 14:34:32 | gibi | yeah nova-multicell did not run on the tempest patch | |
| 14:34:41 | sean-k-mooney | ok so we have two thing one we should skip this temporaly on the multi-cell-job | |
| 14:34:44 | sean-k-mooney | and then fix the bug | |
| 14:34:56 | gibi | yes | |
| 14:34:58 | dansmith | a normal run should have the api pointing at cell0: https://storage.bhs.cloud.ovh.net/v1/AUTH_dcaab5e32b234d56b626f72581e3644c/zuul_opendev_logs_8cf/831219/22/check/tempest-integrated-compute/8cfb267/controller/logs/etc/nova/nova_conf.txt | |
| 14:35:03 | dansmith | but that multicell run has nothing defined there | |
| 14:35:10 | sean-k-mooney | we at least shoudl check we are not unshelving across cells for now but that actully should work | |
| 14:35:18 | sean-k-mooney | since we use shelve for | |
| 14:35:21 | sean-k-mooney | cross cell resize | |
| 14:35:28 | dansmith | sean-k-mooney: it shouldn't work | |
| 14:35:38 | dansmith | because cross-cell migration does other stuff on top of the shelve | |
| 14:35:53 | sean-k-mooney | oh your right | |
| 14:36:22 | gmann | and cross-cell is not in that feature acope right? which is added in this cycle only. | |
| 14:36:24 | gmann | scope | |
| 14:36:33 | gmann | sean-k-mooney: agree to skip it in multi-cell job | |
| 14:36:38 | sean-k-mooney | it was not explcitly no | |
| 14:36:56 | sean-k-mooney | so we can just document the limitation i guess | |
| 14:37:05 | gmann | but will be good to add in documnt/releasenotes | |
| 14:37:05 | gmann | yes | |
| 14:38:07 | sean-k-mooney | dansmith: ohter then usign multiple port binding to test if the destination cell can bind the port. how much more does cross cell migration do over shelve. there is a bunch of stuff in the supper conductor to copy the instnace object between the cell db right | |
| 14:38:33 | dansmith | yeah it moves the thing between DBs, that's the big thing | |
| 14:39:13 | sean-k-mooney | well for now i guess its fine the error could be improved but the api wont let you currpt the db or anything like that | |
| 14:39:27 | dansmith | so this multicell job has two cells and two computes, one compute per cell? it must be disabling live migration and lots of other stuff that would normally require multinode right/ | |
| 14:40:04 | sean-k-mooney | ya its two nodes (contoller and compute) and both nodes are in a differnt cell | |
| 14:40:19 | dansmith | so it must disable a bunch of things | |
| 14:40:31 | sean-k-mooney | yep | |
| 14:41:20 | sean-k-mooney | https://github.com/openstack/nova/blob/master/.zuul.yaml#L514-L571 | |
| 14:41:21 | gmann | yeah, live migration is disabled https://github.com/openstack/nova/blob/master/.zuul.yaml#L554 | |
| 14:41:26 | sean-k-mooney | thats the deffintion | |
| 14:42:36 | dansmith | okay | |
| 14:42:51 | gibi | file the bug https://bugs.launchpad.net/nova/+bug/1988316 | |
| 14:43:05 | gibi | as it is blocks the gate I will set it to Critical | |
| 14:43:17 | dansmith | also, the api config I was looking at was on the compute, the controller is pointed at cell0 by default | |