| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-08-31 | |||
| 14:06:53 | sean-k-mooney | ubuntu-focal-rax-dfw-0030919243 is the compute buntu-focal-rax-dfw-0030919238 is the contoller | |
| 14:07:08 | gmann | this test was recently merged yesterday and was passing that time | |
| 14:07:26 | gibi | it is interesting as the test looks up the other compute by the service list | |
| 14:07:31 | dansmith | might be flaky or might not work on some cloud providers if there's a name weirdness? | |
| 14:07:33 | sean-k-mooney | dansmith: so the name is correct at least it trying to unshelve to the contoller | |
| 14:08:08 | sean-k-mooney | i wonder if there is fqdn stuff going on | |
| 14:08:10 | gibi | ubuntu-focal-rax-dfw-0030919238 doesnt feel wierd | |
| 14:08:23 | gmann | yeah | |
| 14:08:47 | dansmith | gibi: I meant FQDN type things | |
| 14:08:54 | dansmith | sean-k-mooney and I are scarred for life on FQDN problems :) | |
| 14:09:11 | sean-k-mooney | ubuntu-focal-rax-dfw-0030919238 is what we see for the host value in the compute agent startup | |
| 14:09:22 | sean-k-mooney | so maybe hypervior hostname | |
| 14:10:02 | sean-k-mooney | DEBUG nova.compute.resource_tracker [None req-49d84742-3abf-4111-8a83-e2c5e0e67661 None None] Hypervisor/Node resource view: name=ubuntu-focal-rax-dfw-0030919238 | |
| 14:10:18 | sean-k-mooney | nope they all seem to line up | |
| 14:10:43 | sean-k-mooney | odd | |
| 14:11:07 | dansmith | do we do a disabled or service liveness check before we let you unshelve there? maybe the compute is stuck and not updating its counter? | |
| 14:12:15 | gibi | when we query the service we check for that it is up and enabled | |
| 14:12:21 | gmann | dansmith: you mean just before unshelve or while selecting the host? while selecting host we do check service is up and enable | |
| 14:12:43 | gmann | but after that shelve happen and then unshelve and that time no check before unshelve | |
| 14:13:02 | dansmith | gmann: yeah I mean whatever check we're doing that gives that error | |
| 14:13:15 | dansmith | is there a tempest test that is disabling a compute by chance for testing? | |
| 14:13:18 | sean-k-mooney | its proably worth checkign the schduler logs | |
| 14:13:33 | dansmith | oh is this just the scheduler "is this host okay" check? | |
| 14:15:20 | gmann | I do not think we have any such test of touching the compute services enable/disable | |
| 14:15:30 | gibi | it is not about sevice state | |
| 14:15:39 | gibi | the unshelve fails on compute_node_get_all_by_host | |
| 14:15:48 | gibi | so the compute is not in the DBV | |
| 14:15:49 | gibi | DB | |
| 14:17:15 | gibi | https://github.com/openstack/nova/blob/master/nova/compute/api.py#L4577 | |
| 14:17:21 | gibi | this where the unshelve fails | |
| 14:18:53 | sean-k-mooney | so it is acutlly trying to shelve and unshlve multiple times | |
| 14:19:15 | sean-k-mooney | it shleve offloaded and shelve to the compute | |
| 14:19:18 | sean-k-mooney | then it shelved again | |
| 14:19:28 | sean-k-mooney | and failed to unshleve to the contoler with the not found issue | |
| 14:19:52 | gibi | yes | |
| 14:20:12 | gibi | first it unselves back to where it was, then it unshelves to the other compute | |
| 14:20:42 | sean-k-mooney | and req-b3e7705d-a20b-42b3-8eb5-594181a096c8 is the request that fialed | |
| 14:20:52 | sean-k-mooney | which failed in the api right | |
| 14:20:57 | sean-k-mooney | its not in the secheuler | |
| 14:21:02 | gibi | yes | |
| 14:21:07 | gmann | yes it is failing in 2nd time unshelve on another host | |
| 14:21:10 | gibi | it is not fails on the scheduler | |
| 14:21:16 | gibi | it fails in the compute api | |
| 14:21:17 | gibi | https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/controller/logs/screen-n-api.txt | |
| 14:21:20 | gibi | https://github.com/openstack/nova/blob/master/nova/compute/api.py#L4577 | |
| 14:21:33 | gibi | unshelve checks is the requested host exists | |
| 14:21:49 | gibi | before it goes to the conductor | |
| 14:21:56 | sean-k-mooney | by calling compute_node_get_all_by_host | |
| 14:22:20 | gibi | it calls get_first_node_by_host_for_old_compat | |
| 14:22:26 | sean-k-mooney | thats on the main db i.e. cell db right | |
| 14:22:39 | dansmith | so that is looking up by the service, and getting the first compute node on the service | |
| 14:22:45 | dansmith | not by_host_and_node | |
| 14:23:13 | sean-k-mooney | for libvirt there should only be one but ya | |
| 14:23:32 | sean-k-mooney | if shelve was a thing for vmware or ironic that would not be ideal | |
| 14:23:38 | dansmith | sean-k-mooney: right but this is very loose association which doesn't have FK constraints | |
| 14:24:05 | sean-k-mooney | why is it hitting the cell db. it could look at the host mapping table in the api db right | |
| 14:24:14 | sean-k-mooney | or cell mappings table | |
| 14:24:21 | dansmith | eh? | |
| 14:24:29 | dansmith | those just tell us which db the host is in | |
| 14:24:43 | sean-k-mooney | right but if all you care about is does it exist is that not enough | |
| 14:24:51 | dansmith | nothing about whether it's really there or up or whatever | |
| 14:24:51 | gibi | sean-k-mooney: you have a point, so the compute-api does a cell DB query without targeting the cell first | |
| 14:25:23 | sean-k-mooney | dansmith: i guess it could be down but im not sure this is checkign that | |
| 14:25:28 | dansmith | sean-k-mooney: or deleted | |
| 14:25:48 | sean-k-mooney | i would expct that to be check by the schduler to be honest | |
| 14:25:53 | dansmith | gibi: these hosts are in the same cell though so if that wasn't working, we wouldn't have found the first one right? | |
| 14:26:22 | gibi | this is the multicell job so I assume we have hosts in different cells | |
| 14:26:45 | sean-k-mooney | yes there are only two nodes in the job | |
| 14:26:51 | sean-k-mooney | so each is in a different cell | |
| 14:26:59 | dansmith | multicell meaning superconductor not actually one host per cell right? | |
| 14:27:27 | sean-k-mooney | well we can check but i tought this was 2 cells | |
| 14:27:28 | dansmith | or is this really unshelving to different cells? | |
| 14:27:50 | dansmith | can we unshelve across cells? | |
| 14:27:52 | dansmith | I thought not | |
| 14:28:16 | dansmith | so maybe we're targeted at the cell where the instance is and not the cell where the host is that we're unshelving to? | |
| 14:28:24 | dansmith | not sure how that would have ever passed | |
| 14:28:48 | sean-k-mooney | the compute is not runnign a second conductor | |
| 14:28:55 | gibi | I'm looking at the compute.api.API.unshelve and I think the context is not targeted | |
| 14:29:14 | sean-k-mooney | oh the contoler is ruuning both cell1 and cell2 conductors | |
| 14:29:36 | sean-k-mooney | and the supper conductor | |
| 14:30:24 | sean-k-mooney | https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/compute1/logs/etc/nova/nova_cell1_conf.txt | |
| 14:30:51 | sean-k-mooney | transport_url = rabbit://stackrabbit:secretrabbit@10.208.192.130:5672/nova_cell2 | |
| 14:31:00 | sean-k-mooney | so ya compute1 is nova_cell2 | |
| 14:31:20 | sean-k-mooney | and contoler is transport_url = rabbit://stackrabbit:secretrabbit@10.208.192.130:5672/nova_cell1 | |
| 14:31:26 | gibi | I don't know why this ever passed the test | |
| 14:31:45 | dansmith | I also don't know how unshelve could be using an untargeted context | |
| 14:31:54 | dansmith | it should be pointing at cell0 all the time which would never work | |
| 14:32:32 | sean-k-mooney | dansmith: did you say this just started failing recently? | |
| 14:32:42 | sean-k-mooney | because we merged unshelve to host a while ago | |
| 14:32:43 | dansmith | sean-k-mooney: it just recently merged I think gmann said | |
| 14:32:50 | sean-k-mooney | the tempest test? | |
| 14:33:16 | sean-k-mooney | if so i guess we dont run the multi-cell job on tempest | |
| 14:33:44 | gmann | dansmith: sean-k-mooney gibi that passed in tempest-multinode-full-py3 job and I do not think we checked multi-cell job run for this | |
| 14:33:56 | dansmith | the base nova conf doesn't have a db connection pointing at cell0, which I assume is what the apis are using | |
| 14:34:02 | gmann | yeah we do not run multi-cell job there and only checked tempest-multinode-full-py3 job passing it | |
| 14:34:22 | sean-k-mooney | gmann: ack that makes sense | |
| 14:34:32 | gibi | yeah nova-multicell did not run on the tempest patch | |
| 14:34:41 | sean-k-mooney | ok so we have two thing one we should skip this temporaly on the multi-cell-job | |
| 14:34:44 | sean-k-mooney | and then fix the bug | |
| 14:34:56 | gibi | yes | |