Earlier  
Posted Nick Remark
#openstack-nova - 2022-08-31
12:39:07 sean-k-mooney if gibi or bauzas dont do it before
12:39:22 gibi I'm on it
12:40:11 whoami-rajat thanks
13:42:12 dansmith gibi: wanna hear something funny?
13:51:40 gibi dansmith: sure
13:52:09 dansmith gibi: I would have bet my lunch on user_data being in the set of things we don't query out by default and only load if required
13:52:23 dansmith I remember extensive discussions about it around icehouse when all this was being done
13:52:30 dansmith because of it's size
13:52:49 dansmith but I think we're actually *always* pulling that out and *always* sending it over the bus right now, which is insane because it's almost never used
13:53:21 dansmith *and* in any of the joined queries where we duplicate the instance fields (like why we stopped joining the metadata queries), we'd be duplicating up to 64k of user data in the DB response...
13:53:59 sean-k-mooney dansmith: we had a cve related to that in the past
13:54:10 sean-k-mooney that should be fixed now
13:54:16 dansmith sean-k-mooney: related to what?
13:54:33 gibi I don't see any special casing for user_data so probably you are right we are always pulling it out
13:54:35 sean-k-mooney pulling GBs of results into a join when you have large userdata
13:54:52 dansmith okay, then fixed how?
13:55:04 dansmith because AFAICT, we're always pulling it
13:55:07 sean-k-mooney https://review.opendev.org/c/openstack/nova/+/758928/
13:55:49 dansmith sean-k-mooney: that's metadata
13:56:00 dansmith but you're saying that because we were also loading user_data that was much larger I guess?
13:56:24 sean-k-mooney yep so i think we fixed the once case wehre we had a really large join as a result
13:56:53 dansmith yeah so that patch will maybe avoid some duplication of it, but we don't need to be pulling it out of the DB except when we need it
13:57:14 sean-k-mooney yes thats also fair we dont
13:58:25 sean-k-mooney i guess it would be nice to do as a performace enhancement in general
13:58:34 gibi does it also mean that we can save the updated user data on the compute side now?
13:58:37 sean-k-mooney but also possibly reducing the change of another cve
13:58:38 dansmith well, it could seriously cut down on our rabbit load
13:58:57 sean-k-mooney where the user data is large yes
13:59:04 dansmith gibi: yeah I assume that if that patch works, it works because somewhere in reboot the instance gets saved, probably due to state
13:59:15 dansmith gibi: no test coverage of that though that I can see :)
13:59:17 sean-k-mooney unfortunetly people that use heat tend to also abuse the userdata
13:59:29 sean-k-mooney we have had requests in teh past to increae the limmit even more
13:59:32 gibi dansmith: ahh yes as we dont call save on the api side
13:59:37 sean-k-mooney we said no the last few times it came up
13:59:58 dansmith sean-k-mooney: yeah, I bet people are bzipping stuff to keep it under the limit :D
14:01:27 sean-k-mooney oh by the way the limit of medium text is 64MB not 64KB but i think we limit to 64KB but now i want to go check
14:01:41 sean-k-mooney we use mediumtext in the db schema
14:01:46 sean-k-mooney i think we put the limit in the api
14:01:59 dansmith we are limiting in the api yeah
14:02:06 dansmith at least in this patch
14:02:58 sean-k-mooney our api ref say Configuration information or scripts to use upon launch. Must be Base64 encoded. Restricted to 65535 bytes.
14:03:12 sean-k-mooney which is what i expected
14:03:25 dansmith btw, I've seen this fail several times in the last day: https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/testr_results.html
14:04:14 dansmith it complains of a missing host, which the test is unshelving to by name, so I assume that's a test bug or something
14:04:16 sean-k-mooney Compute host ubuntu-focal-rax-dfw-0030919238 could not be found.
14:04:20 sean-k-mooney ya
14:04:28 sean-k-mooney so tempest config error maybe
14:05:15 dansmith so anyway, on the user_data thing, I think it would be a good idea for us to en-lazy that by default and only load it in the api and metadata api by default and I bet we'll see some rabbit load relaxed
14:06:15 gmann dansmith: sean-k-mooney gibi Uggla I am also seeing nova-multi-cell failing consistently for the new test added for test_unshelve_to_specific_host
14:06:36 gmann https://zuul.opendev.org/t/openstack/builds?job_name=nova-multi-cell&skip=0
14:06:45 dansmith ack, I've probably rechecked 8 of that failure in the last 24 hours
14:06:53 sean-k-mooney ubuntu-focal-rax-dfw-0030919243 is the compute buntu-focal-rax-dfw-0030919238 is the contoller
14:07:08 gmann this test was recently merged yesterday and was passing that time
14:07:26 gibi it is interesting as the test looks up the other compute by the service list
14:07:31 dansmith might be flaky or might not work on some cloud providers if there's a name weirdness?
14:07:33 sean-k-mooney dansmith: so the name is correct at least it trying to unshelve to the contoller
14:08:08 sean-k-mooney i wonder if there is fqdn stuff going on
14:08:10 gibi ubuntu-focal-rax-dfw-0030919238 doesnt feel wierd
14:08:23 gmann yeah
14:08:47 dansmith gibi: I meant FQDN type things
14:08:54 dansmith sean-k-mooney and I are scarred for life on FQDN problems :)
14:09:11 sean-k-mooney ubuntu-focal-rax-dfw-0030919238 is what we see for the host value in the compute agent startup
14:09:22 sean-k-mooney so maybe hypervior hostname
14:10:02 sean-k-mooney DEBUG nova.compute.resource_tracker [None req-49d84742-3abf-4111-8a83-e2c5e0e67661 None None] Hypervisor/Node resource view: name=ubuntu-focal-rax-dfw-0030919238
14:10:18 sean-k-mooney nope they all seem to line up
14:10:43 sean-k-mooney odd
14:11:07 dansmith do we do a disabled or service liveness check before we let you unshelve there? maybe the compute is stuck and not updating its counter?
14:12:15 gibi when we query the service we check for that it is up and enabled
14:12:21 gmann dansmith: you mean just before unshelve or while selecting the host? while selecting host we do check service is up and enable
14:12:43 gmann but after that shelve happen and then unshelve and that time no check before unshelve
14:13:02 dansmith gmann: yeah I mean whatever check we're doing that gives that error
14:13:15 dansmith is there a tempest test that is disabling a compute by chance for testing?
14:13:18 sean-k-mooney its proably worth checkign the schduler logs
14:13:33 dansmith oh is this just the scheduler "is this host okay" check?
14:15:20 gmann I do not think we have any such test of touching the compute services enable/disable
14:15:30 gibi it is not about sevice state
14:15:39 gibi the unshelve fails on compute_node_get_all_by_host
14:15:48 gibi so the compute is not in the DBV
14:15:49 gibi DB
14:17:15 gibi https://github.com/openstack/nova/blob/master/nova/compute/api.py#L4577
14:17:21 gibi this where the unshelve fails
14:18:53 sean-k-mooney so it is acutlly trying to shelve and unshlve multiple times
14:19:15 sean-k-mooney it shleve offloaded and shelve to the compute
14:19:18 sean-k-mooney then it shelved again
14:19:28 sean-k-mooney and failed to unshleve to the contoler with the not found issue
14:19:52 gibi yes
14:20:12 gibi first it unselves back to where it was, then it unshelves to the other compute
14:20:42 sean-k-mooney and req-b3e7705d-a20b-42b3-8eb5-594181a096c8 is the request that fialed
14:20:52 sean-k-mooney which failed in the api right
14:20:57 sean-k-mooney its not in the secheuler
14:21:02 gibi yes
14:21:07 gmann yes it is failing in 2nd time unshelve on another host
14:21:10 gibi it is not fails on the scheduler
14:21:16 gibi it fails in the compute api
14:21:17 gibi https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/controller/logs/screen-n-api.txt
14:21:20 gibi https://github.com/openstack/nova/blob/master/nova/compute/api.py#L4577
14:21:33 gibi unshelve checks is the requested host exists
14:21:49 gibi before it goes to the conductor
14:21:56 sean-k-mooney by calling compute_node_get_all_by_host
14:22:20 gibi it calls get_first_node_by_host_for_old_compat

Earlier   Later