| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-08-31 | |||
| 12:10:51 | whoami-rajat | sean-k-mooney, np, i had to rebase so was just avoiding followups | |
| 12:11:10 | sean-k-mooney | whoami-rajat: and sorry for the micorversion hassel. stacking the changes was ment to prevent fighting for the same microversion but in this case it did not help | |
| 12:12:08 | opendevreview | Takashi Natsume proposed openstack/nova master: Add missing descriptions in HACKING.rst https://review.opendev.org/c/openstack/nova/+/853054 | |
| 12:12:18 | opendevreview | Takashi Natsume proposed openstack/nova master: Add a hacking rule for the setDaemon method https://review.opendev.org/c/openstack/nova/+/854653 | |
| 12:12:27 | whoami-rajat | sean-k-mooney, i can understand there were some last minute decisions to be made, everything is fine until the changes get in so no problem :) | |
| 12:36:35 | opendevreview | Rajat Dhasmana proposed openstack/nova master: Add API support for rebuilding BFV instances https://review.opendev.org/c/openstack/nova/+/830883 | |
| 12:37:34 | whoami-rajat | ^ there was a functional test failure due to the method name change in compute manager, everything else should is running fine | |
| 12:38:56 | sean-k-mooney | ack ill re reivew after donwstream call | |
| 12:39:07 | sean-k-mooney | if gibi or bauzas dont do it before | |
| 12:39:22 | gibi | I'm on it | |
| 12:40:11 | whoami-rajat | thanks | |
| 13:42:12 | dansmith | gibi: wanna hear something funny? | |
| 13:51:40 | gibi | dansmith: sure | |
| 13:52:09 | dansmith | gibi: I would have bet my lunch on user_data being in the set of things we don't query out by default and only load if required | |
| 13:52:23 | dansmith | I remember extensive discussions about it around icehouse when all this was being done | |
| 13:52:30 | dansmith | because of it's size | |
| 13:52:49 | dansmith | but I think we're actually *always* pulling that out and *always* sending it over the bus right now, which is insane because it's almost never used | |
| 13:53:21 | dansmith | *and* in any of the joined queries where we duplicate the instance fields (like why we stopped joining the metadata queries), we'd be duplicating up to 64k of user data in the DB response... | |
| 13:53:59 | sean-k-mooney | dansmith: we had a cve related to that in the past | |
| 13:54:10 | sean-k-mooney | that should be fixed now | |
| 13:54:16 | dansmith | sean-k-mooney: related to what? | |
| 13:54:33 | gibi | I don't see any special casing for user_data so probably you are right we are always pulling it out | |
| 13:54:35 | sean-k-mooney | pulling GBs of results into a join when you have large userdata | |
| 13:54:52 | dansmith | okay, then fixed how? | |
| 13:55:04 | dansmith | because AFAICT, we're always pulling it | |
| 13:55:07 | sean-k-mooney | https://review.opendev.org/c/openstack/nova/+/758928/ | |
| 13:55:49 | dansmith | sean-k-mooney: that's metadata | |
| 13:56:00 | dansmith | but you're saying that because we were also loading user_data that was much larger I guess? | |
| 13:56:24 | sean-k-mooney | yep so i think we fixed the once case wehre we had a really large join as a result | |
| 13:56:53 | dansmith | yeah so that patch will maybe avoid some duplication of it, but we don't need to be pulling it out of the DB except when we need it | |
| 13:57:14 | sean-k-mooney | yes thats also fair we dont | |
| 13:58:25 | sean-k-mooney | i guess it would be nice to do as a performace enhancement in general | |
| 13:58:34 | gibi | does it also mean that we can save the updated user data on the compute side now? | |
| 13:58:37 | sean-k-mooney | but also possibly reducing the change of another cve | |
| 13:58:38 | dansmith | well, it could seriously cut down on our rabbit load | |
| 13:58:57 | sean-k-mooney | where the user data is large yes | |
| 13:59:04 | dansmith | gibi: yeah I assume that if that patch works, it works because somewhere in reboot the instance gets saved, probably due to state | |
| 13:59:15 | dansmith | gibi: no test coverage of that though that I can see :) | |
| 13:59:17 | sean-k-mooney | unfortunetly people that use heat tend to also abuse the userdata | |
| 13:59:29 | sean-k-mooney | we have had requests in teh past to increae the limmit even more | |
| 13:59:32 | gibi | dansmith: ahh yes as we dont call save on the api side | |
| 13:59:37 | sean-k-mooney | we said no the last few times it came up | |
| 13:59:58 | dansmith | sean-k-mooney: yeah, I bet people are bzipping stuff to keep it under the limit :D | |
| 14:01:27 | sean-k-mooney | oh by the way the limit of medium text is 64MB not 64KB but i think we limit to 64KB but now i want to go check | |
| 14:01:41 | sean-k-mooney | we use mediumtext in the db schema | |
| 14:01:46 | sean-k-mooney | i think we put the limit in the api | |
| 14:01:59 | dansmith | we are limiting in the api yeah | |
| 14:02:06 | dansmith | at least in this patch | |
| 14:02:58 | sean-k-mooney | our api ref say Configuration information or scripts to use upon launch. Must be Base64 encoded. Restricted to 65535 bytes. | |
| 14:03:12 | sean-k-mooney | which is what i expected | |
| 14:03:25 | dansmith | btw, I've seen this fail several times in the last day: https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/testr_results.html | |
| 14:04:14 | dansmith | it complains of a missing host, which the test is unshelving to by name, so I assume that's a test bug or something | |
| 14:04:16 | sean-k-mooney | Compute host ubuntu-focal-rax-dfw-0030919238 could not be found. | |
| 14:04:20 | sean-k-mooney | ya | |
| 14:04:28 | sean-k-mooney | so tempest config error maybe | |
| 14:05:15 | dansmith | so anyway, on the user_data thing, I think it would be a good idea for us to en-lazy that by default and only load it in the api and metadata api by default and I bet we'll see some rabbit load relaxed | |
| 14:06:15 | gmann | dansmith: sean-k-mooney gibi Uggla I am also seeing nova-multi-cell failing consistently for the new test added for test_unshelve_to_specific_host | |
| 14:06:36 | gmann | https://zuul.opendev.org/t/openstack/builds?job_name=nova-multi-cell&skip=0 | |
| 14:06:45 | dansmith | ack, I've probably rechecked 8 of that failure in the last 24 hours | |
| 14:06:53 | sean-k-mooney | ubuntu-focal-rax-dfw-0030919243 is the compute buntu-focal-rax-dfw-0030919238 is the contoller | |
| 14:07:08 | gmann | this test was recently merged yesterday and was passing that time | |
| 14:07:26 | gibi | it is interesting as the test looks up the other compute by the service list | |
| 14:07:31 | dansmith | might be flaky or might not work on some cloud providers if there's a name weirdness? | |
| 14:07:33 | sean-k-mooney | dansmith: so the name is correct at least it trying to unshelve to the contoller | |
| 14:08:08 | sean-k-mooney | i wonder if there is fqdn stuff going on | |
| 14:08:10 | gibi | ubuntu-focal-rax-dfw-0030919238 doesnt feel wierd | |
| 14:08:23 | gmann | yeah | |
| 14:08:47 | dansmith | gibi: I meant FQDN type things | |
| 14:08:54 | dansmith | sean-k-mooney and I are scarred for life on FQDN problems :) | |
| 14:09:11 | sean-k-mooney | ubuntu-focal-rax-dfw-0030919238 is what we see for the host value in the compute agent startup | |
| 14:09:22 | sean-k-mooney | so maybe hypervior hostname | |
| 14:10:02 | sean-k-mooney | DEBUG nova.compute.resource_tracker [None req-49d84742-3abf-4111-8a83-e2c5e0e67661 None None] Hypervisor/Node resource view: name=ubuntu-focal-rax-dfw-0030919238 | |
| 14:10:18 | sean-k-mooney | nope they all seem to line up | |
| 14:10:43 | sean-k-mooney | odd | |
| 14:11:07 | dansmith | do we do a disabled or service liveness check before we let you unshelve there? maybe the compute is stuck and not updating its counter? | |
| 14:12:15 | gibi | when we query the service we check for that it is up and enabled | |
| 14:12:21 | gmann | dansmith: you mean just before unshelve or while selecting the host? while selecting host we do check service is up and enable | |
| 14:12:43 | gmann | but after that shelve happen and then unshelve and that time no check before unshelve | |
| 14:13:02 | dansmith | gmann: yeah I mean whatever check we're doing that gives that error | |
| 14:13:15 | dansmith | is there a tempest test that is disabling a compute by chance for testing? | |
| 14:13:18 | sean-k-mooney | its proably worth checkign the schduler logs | |
| 14:13:33 | dansmith | oh is this just the scheduler "is this host okay" check? | |
| 14:15:20 | gmann | I do not think we have any such test of touching the compute services enable/disable | |
| 14:15:30 | gibi | it is not about sevice state | |
| 14:15:39 | gibi | the unshelve fails on compute_node_get_all_by_host | |
| 14:15:48 | gibi | so the compute is not in the DBV | |
| 14:15:49 | gibi | DB | |
| 14:17:15 | gibi | https://github.com/openstack/nova/blob/master/nova/compute/api.py#L4577 | |
| 14:17:21 | gibi | this where the unshelve fails | |
| 14:18:53 | sean-k-mooney | so it is acutlly trying to shelve and unshlve multiple times | |
| 14:19:15 | sean-k-mooney | it shleve offloaded and shelve to the compute | |
| 14:19:18 | sean-k-mooney | then it shelved again | |
| 14:19:28 | sean-k-mooney | and failed to unshleve to the contoler with the not found issue | |
| 14:19:52 | gibi | yes | |
| 14:20:12 | gibi | first it unselves back to where it was, then it unshelves to the other compute | |
| 14:20:42 | sean-k-mooney | and req-b3e7705d-a20b-42b3-8eb5-594181a096c8 is the request that fialed | |
| 14:20:52 | sean-k-mooney | which failed in the api right | |
| 14:20:57 | sean-k-mooney | its not in the secheuler | |
| 14:21:02 | gibi | yes | |
| 14:21:07 | gmann | yes it is failing in 2nd time unshelve on another host | |