Earlier  
Posted Nick Remark
#openstack-nova - 2022-08-31
14:02:58 sean-k-mooney our api ref say Configuration information or scripts to use upon launch. Must be Base64 encoded. Restricted to 65535 bytes.
14:03:12 sean-k-mooney which is what i expected
14:03:25 dansmith btw, I've seen this fail several times in the last day: https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/testr_results.html
14:04:14 dansmith it complains of a missing host, which the test is unshelving to by name, so I assume that's a test bug or something
14:04:16 sean-k-mooney Compute host ubuntu-focal-rax-dfw-0030919238 could not be found.
14:04:20 sean-k-mooney ya
14:04:28 sean-k-mooney so tempest config error maybe
14:05:15 dansmith so anyway, on the user_data thing, I think it would be a good idea for us to en-lazy that by default and only load it in the api and metadata api by default and I bet we'll see some rabbit load relaxed
14:06:15 gmann dansmith: sean-k-mooney gibi Uggla I am also seeing nova-multi-cell failing consistently for the new test added for test_unshelve_to_specific_host
14:06:36 gmann https://zuul.opendev.org/t/openstack/builds?job_name=nova-multi-cell&skip=0
14:06:45 dansmith ack, I've probably rechecked 8 of that failure in the last 24 hours
14:06:53 sean-k-mooney ubuntu-focal-rax-dfw-0030919243 is the compute buntu-focal-rax-dfw-0030919238 is the contoller
14:07:08 gmann this test was recently merged yesterday and was passing that time
14:07:26 gibi it is interesting as the test looks up the other compute by the service list
14:07:31 dansmith might be flaky or might not work on some cloud providers if there's a name weirdness?
14:07:33 sean-k-mooney dansmith: so the name is correct at least it trying to unshelve to the contoller
14:08:08 sean-k-mooney i wonder if there is fqdn stuff going on
14:08:10 gibi ubuntu-focal-rax-dfw-0030919238 doesnt feel wierd
14:08:23 gmann yeah
14:08:47 dansmith gibi: I meant FQDN type things
14:08:54 dansmith sean-k-mooney and I are scarred for life on FQDN problems :)
14:09:11 sean-k-mooney ubuntu-focal-rax-dfw-0030919238 is what we see for the host value in the compute agent startup
14:09:22 sean-k-mooney so maybe hypervior hostname
14:10:02 sean-k-mooney DEBUG nova.compute.resource_tracker [None req-49d84742-3abf-4111-8a83-e2c5e0e67661 None None] Hypervisor/Node resource view: name=ubuntu-focal-rax-dfw-0030919238
14:10:18 sean-k-mooney nope they all seem to line up
14:10:43 sean-k-mooney odd
14:11:07 dansmith do we do a disabled or service liveness check before we let you unshelve there? maybe the compute is stuck and not updating its counter?
14:12:15 gibi when we query the service we check for that it is up and enabled
14:12:21 gmann dansmith: you mean just before unshelve or while selecting the host? while selecting host we do check service is up and enable
14:12:43 gmann but after that shelve happen and then unshelve and that time no check before unshelve
14:13:02 dansmith gmann: yeah I mean whatever check we're doing that gives that error
14:13:15 dansmith is there a tempest test that is disabling a compute by chance for testing?
14:13:18 sean-k-mooney its proably worth checkign the schduler logs
14:13:33 dansmith oh is this just the scheduler "is this host okay" check?
14:15:20 gmann I do not think we have any such test of touching the compute services enable/disable
14:15:30 gibi it is not about sevice state
14:15:39 gibi the unshelve fails on compute_node_get_all_by_host
14:15:48 gibi so the compute is not in the DBV
14:15:49 gibi DB
14:17:15 gibi https://github.com/openstack/nova/blob/master/nova/compute/api.py#L4577
14:17:21 gibi this where the unshelve fails
14:18:53 sean-k-mooney so it is acutlly trying to shelve and unshlve multiple times
14:19:15 sean-k-mooney it shleve offloaded and shelve to the compute
14:19:18 sean-k-mooney then it shelved again
14:19:28 sean-k-mooney and failed to unshleve to the contoler with the not found issue
14:19:52 gibi yes
14:20:12 gibi first it unselves back to where it was, then it unshelves to the other compute
14:20:42 sean-k-mooney and req-b3e7705d-a20b-42b3-8eb5-594181a096c8 is the request that fialed
14:20:52 sean-k-mooney which failed in the api right
14:20:57 sean-k-mooney its not in the secheuler
14:21:02 gibi yes
14:21:07 gmann yes it is failing in 2nd time unshelve on another host
14:21:10 gibi it is not fails on the scheduler
14:21:16 gibi it fails in the compute api
14:21:17 gibi https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/controller/logs/screen-n-api.txt
14:21:20 gibi https://github.com/openstack/nova/blob/master/nova/compute/api.py#L4577
14:21:33 gibi unshelve checks is the requested host exists
14:21:49 gibi before it goes to the conductor
14:21:56 sean-k-mooney by calling compute_node_get_all_by_host
14:22:20 gibi it calls get_first_node_by_host_for_old_compat
14:22:26 sean-k-mooney thats on the main db i.e. cell db right
14:22:39 dansmith so that is looking up by the service, and getting the first compute node on the service
14:22:45 dansmith not by_host_and_node
14:23:13 sean-k-mooney for libvirt there should only be one but ya
14:23:32 sean-k-mooney if shelve was a thing for vmware or ironic that would not be ideal
14:23:38 dansmith sean-k-mooney: right but this is very loose association which doesn't have FK constraints
14:24:05 sean-k-mooney why is it hitting the cell db. it could look at the host mapping table in the api db right
14:24:14 sean-k-mooney or cell mappings table
14:24:21 dansmith eh?
14:24:29 dansmith those just tell us which db the host is in
14:24:43 sean-k-mooney right but if all you care about is does it exist is that not enough
14:24:51 gibi sean-k-mooney: you have a point, so the compute-api does a cell DB query without targeting the cell first
14:24:51 dansmith nothing about whether it's really there or up or whatever
14:25:23 sean-k-mooney dansmith: i guess it could be down but im not sure this is checkign that
14:25:28 dansmith sean-k-mooney: or deleted
14:25:48 sean-k-mooney i would expct that to be check by the schduler to be honest
14:25:53 dansmith gibi: these hosts are in the same cell though so if that wasn't working, we wouldn't have found the first one right?
14:26:22 gibi this is the multicell job so I assume we have hosts in different cells
14:26:45 sean-k-mooney yes there are only two nodes in the job
14:26:51 sean-k-mooney so each is in a different cell
14:26:59 dansmith multicell meaning superconductor not actually one host per cell right?
14:27:27 sean-k-mooney well we can check but i tought this was 2 cells
14:27:28 dansmith or is this really unshelving to different cells?
14:27:50 dansmith can we unshelve across cells?
14:27:52 dansmith I thought not
14:28:16 dansmith so maybe we're targeted at the cell where the instance is and not the cell where the host is that we're unshelving to?
14:28:24 dansmith not sure how that would have ever passed
14:28:48 sean-k-mooney the compute is not runnign a second conductor
14:28:55 gibi I'm looking at the compute.api.API.unshelve and I think the context is not targeted
14:29:14 sean-k-mooney oh the contoler is ruuning both cell1 and cell2 conductors
14:29:36 sean-k-mooney and the supper conductor
14:30:24 sean-k-mooney https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/compute1/logs/etc/nova/nova_cell1_conf.txt
14:30:51 sean-k-mooney transport_url = rabbit://stackrabbit:secretrabbit@10.208.192.130:5672/nova_cell2
14:31:00 sean-k-mooney so ya compute1 is nova_cell2
14:31:20 sean-k-mooney and contoler is transport_url = rabbit://stackrabbit:secretrabbit@10.208.192.130:5672/nova_cell1
14:31:26 gibi I don't know why this ever passed the test
14:31:45 dansmith I also don't know how unshelve could be using an untargeted context
14:31:54 dansmith it should be pointing at cell0 all the time which would never work
14:32:32 sean-k-mooney dansmith: did you say this just started failing recently?
14:32:42 sean-k-mooney because we merged unshelve to host a while ago

Earlier   Later