Earlier  
Posted Nick Remark
#openstack-nova - 2022-08-31
13:55:07 sean-k-mooney https://review.opendev.org/c/openstack/nova/+/758928/
13:55:49 dansmith sean-k-mooney: that's metadata
13:56:00 dansmith but you're saying that because we were also loading user_data that was much larger I guess?
13:56:24 sean-k-mooney yep so i think we fixed the once case wehre we had a really large join as a result
13:56:53 dansmith yeah so that patch will maybe avoid some duplication of it, but we don't need to be pulling it out of the DB except when we need it
13:57:14 sean-k-mooney yes thats also fair we dont
13:58:25 sean-k-mooney i guess it would be nice to do as a performace enhancement in general
13:58:34 gibi does it also mean that we can save the updated user data on the compute side now?
13:58:37 sean-k-mooney but also possibly reducing the change of another cve
13:58:38 dansmith well, it could seriously cut down on our rabbit load
13:58:57 sean-k-mooney where the user data is large yes
13:59:04 dansmith gibi: yeah I assume that if that patch works, it works because somewhere in reboot the instance gets saved, probably due to state
13:59:15 dansmith gibi: no test coverage of that though that I can see :)
13:59:17 sean-k-mooney unfortunetly people that use heat tend to also abuse the userdata
13:59:29 sean-k-mooney we have had requests in teh past to increae the limmit even more
13:59:32 gibi dansmith: ahh yes as we dont call save on the api side
13:59:37 sean-k-mooney we said no the last few times it came up
13:59:58 dansmith sean-k-mooney: yeah, I bet people are bzipping stuff to keep it under the limit :D
14:01:27 sean-k-mooney oh by the way the limit of medium text is 64MB not 64KB but i think we limit to 64KB but now i want to go check
14:01:41 sean-k-mooney we use mediumtext in the db schema
14:01:46 sean-k-mooney i think we put the limit in the api
14:01:59 dansmith we are limiting in the api yeah
14:02:06 dansmith at least in this patch
14:02:58 sean-k-mooney our api ref say Configuration information or scripts to use upon launch. Must be Base64 encoded. Restricted to 65535 bytes.
14:03:12 sean-k-mooney which is what i expected
14:03:25 dansmith btw, I've seen this fail several times in the last day: https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/testr_results.html
14:04:14 dansmith it complains of a missing host, which the test is unshelving to by name, so I assume that's a test bug or something
14:04:16 sean-k-mooney Compute host ubuntu-focal-rax-dfw-0030919238 could not be found.
14:04:20 sean-k-mooney ya
14:04:28 sean-k-mooney so tempest config error maybe
14:05:15 dansmith so anyway, on the user_data thing, I think it would be a good idea for us to en-lazy that by default and only load it in the api and metadata api by default and I bet we'll see some rabbit load relaxed
14:06:15 gmann dansmith: sean-k-mooney gibi Uggla I am also seeing nova-multi-cell failing consistently for the new test added for test_unshelve_to_specific_host
14:06:36 gmann https://zuul.opendev.org/t/openstack/builds?job_name=nova-multi-cell&skip=0
14:06:45 dansmith ack, I've probably rechecked 8 of that failure in the last 24 hours
14:06:53 sean-k-mooney ubuntu-focal-rax-dfw-0030919243 is the compute buntu-focal-rax-dfw-0030919238 is the contoller
14:07:08 gmann this test was recently merged yesterday and was passing that time
14:07:26 gibi it is interesting as the test looks up the other compute by the service list
14:07:31 dansmith might be flaky or might not work on some cloud providers if there's a name weirdness?
14:07:33 sean-k-mooney dansmith: so the name is correct at least it trying to unshelve to the contoller
14:08:08 sean-k-mooney i wonder if there is fqdn stuff going on
14:08:10 gibi ubuntu-focal-rax-dfw-0030919238 doesnt feel wierd
14:08:23 gmann yeah
14:08:47 dansmith gibi: I meant FQDN type things
14:08:54 dansmith sean-k-mooney and I are scarred for life on FQDN problems :)
14:09:11 sean-k-mooney ubuntu-focal-rax-dfw-0030919238 is what we see for the host value in the compute agent startup
14:09:22 sean-k-mooney so maybe hypervior hostname
14:10:02 sean-k-mooney DEBUG nova.compute.resource_tracker [None req-49d84742-3abf-4111-8a83-e2c5e0e67661 None None] Hypervisor/Node resource view: name=ubuntu-focal-rax-dfw-0030919238
14:10:18 sean-k-mooney nope they all seem to line up
14:10:43 sean-k-mooney odd
14:11:07 dansmith do we do a disabled or service liveness check before we let you unshelve there? maybe the compute is stuck and not updating its counter?
14:12:15 gibi when we query the service we check for that it is up and enabled
14:12:21 gmann dansmith: you mean just before unshelve or while selecting the host? while selecting host we do check service is up and enable
14:12:43 gmann but after that shelve happen and then unshelve and that time no check before unshelve
14:13:02 dansmith gmann: yeah I mean whatever check we're doing that gives that error
14:13:15 dansmith is there a tempest test that is disabling a compute by chance for testing?
14:13:18 sean-k-mooney its proably worth checkign the schduler logs
14:13:33 dansmith oh is this just the scheduler "is this host okay" check?
14:15:20 gmann I do not think we have any such test of touching the compute services enable/disable
14:15:30 gibi it is not about sevice state
14:15:39 gibi the unshelve fails on compute_node_get_all_by_host
14:15:48 gibi so the compute is not in the DBV
14:15:49 gibi DB
14:17:15 gibi https://github.com/openstack/nova/blob/master/nova/compute/api.py#L4577
14:17:21 gibi this where the unshelve fails
14:18:53 sean-k-mooney so it is acutlly trying to shelve and unshlve multiple times
14:19:15 sean-k-mooney it shleve offloaded and shelve to the compute
14:19:18 sean-k-mooney then it shelved again
14:19:28 sean-k-mooney and failed to unshleve to the contoler with the not found issue
14:19:52 gibi yes
14:20:12 gibi first it unselves back to where it was, then it unshelves to the other compute
14:20:42 sean-k-mooney and req-b3e7705d-a20b-42b3-8eb5-594181a096c8 is the request that fialed
14:20:52 sean-k-mooney which failed in the api right
14:20:57 sean-k-mooney its not in the secheuler
14:21:02 gibi yes
14:21:07 gmann yes it is failing in 2nd time unshelve on another host
14:21:10 gibi it is not fails on the scheduler
14:21:16 gibi it fails in the compute api
14:21:17 gibi https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/controller/logs/screen-n-api.txt
14:21:20 gibi https://github.com/openstack/nova/blob/master/nova/compute/api.py#L4577
14:21:33 gibi unshelve checks is the requested host exists
14:21:49 gibi before it goes to the conductor
14:21:56 sean-k-mooney by calling compute_node_get_all_by_host
14:22:20 gibi it calls get_first_node_by_host_for_old_compat
14:22:26 sean-k-mooney thats on the main db i.e. cell db right
14:22:39 dansmith so that is looking up by the service, and getting the first compute node on the service
14:22:45 dansmith not by_host_and_node
14:23:13 sean-k-mooney for libvirt there should only be one but ya
14:23:32 sean-k-mooney if shelve was a thing for vmware or ironic that would not be ideal
14:23:38 dansmith sean-k-mooney: right but this is very loose association which doesn't have FK constraints
14:24:05 sean-k-mooney why is it hitting the cell db. it could look at the host mapping table in the api db right
14:24:14 sean-k-mooney or cell mappings table
14:24:21 dansmith eh?
14:24:29 dansmith those just tell us which db the host is in
14:24:43 sean-k-mooney right but if all you care about is does it exist is that not enough
14:24:51 dansmith nothing about whether it's really there or up or whatever
14:24:51 gibi sean-k-mooney: you have a point, so the compute-api does a cell DB query without targeting the cell first
14:25:23 sean-k-mooney dansmith: i guess it could be down but im not sure this is checkign that
14:25:28 dansmith sean-k-mooney: or deleted
14:25:48 sean-k-mooney i would expct that to be check by the schduler to be honest
14:25:53 dansmith gibi: these hosts are in the same cell though so if that wasn't working, we wouldn't have found the first one right?

Earlier   Later