| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-08-31 | |||
| 14:01:27 | sean-k-mooney | oh by the way the limit of medium text is 64MB not 64KB but i think we limit to 64KB but now i want to go check | |
| 14:01:41 | sean-k-mooney | we use mediumtext in the db schema | |
| 14:01:46 | sean-k-mooney | i think we put the limit in the api | |
| 14:01:59 | dansmith | we are limiting in the api yeah | |
| 14:02:06 | dansmith | at least in this patch | |
| 14:02:58 | sean-k-mooney | our api ref say Configuration information or scripts to use upon launch. Must be Base64 encoded. Restricted to 65535 bytes. | |
| 14:03:12 | sean-k-mooney | which is what i expected | |
| 14:03:25 | dansmith | btw, I've seen this fail several times in the last day: https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/testr_results.html | |
| 14:04:14 | dansmith | it complains of a missing host, which the test is unshelving to by name, so I assume that's a test bug or something | |
| 14:04:16 | sean-k-mooney | Compute host ubuntu-focal-rax-dfw-0030919238 could not be found. | |
| 14:04:20 | sean-k-mooney | ya | |
| 14:04:28 | sean-k-mooney | so tempest config error maybe | |
| 14:05:15 | dansmith | so anyway, on the user_data thing, I think it would be a good idea for us to en-lazy that by default and only load it in the api and metadata api by default and I bet we'll see some rabbit load relaxed | |
| 14:06:15 | gmann | dansmith: sean-k-mooney gibi Uggla I am also seeing nova-multi-cell failing consistently for the new test added for test_unshelve_to_specific_host | |
| 14:06:36 | gmann | https://zuul.opendev.org/t/openstack/builds?job_name=nova-multi-cell&skip=0 | |
| 14:06:45 | dansmith | ack, I've probably rechecked 8 of that failure in the last 24 hours | |
| 14:06:53 | sean-k-mooney | ubuntu-focal-rax-dfw-0030919243 is the compute buntu-focal-rax-dfw-0030919238 is the contoller | |
| 14:07:08 | gmann | this test was recently merged yesterday and was passing that time | |
| 14:07:26 | gibi | it is interesting as the test looks up the other compute by the service list | |
| 14:07:31 | dansmith | might be flaky or might not work on some cloud providers if there's a name weirdness? | |
| 14:07:33 | sean-k-mooney | dansmith: so the name is correct at least it trying to unshelve to the contoller | |
| 14:08:08 | sean-k-mooney | i wonder if there is fqdn stuff going on | |
| 14:08:10 | gibi | ubuntu-focal-rax-dfw-0030919238 doesnt feel wierd | |
| 14:08:23 | gmann | yeah | |
| 14:08:47 | dansmith | gibi: I meant FQDN type things | |
| 14:08:54 | dansmith | sean-k-mooney and I are scarred for life on FQDN problems :) | |
| 14:09:11 | sean-k-mooney | ubuntu-focal-rax-dfw-0030919238 is what we see for the host value in the compute agent startup | |
| 14:09:22 | sean-k-mooney | so maybe hypervior hostname | |
| 14:10:02 | sean-k-mooney | DEBUG nova.compute.resource_tracker [None req-49d84742-3abf-4111-8a83-e2c5e0e67661 None None] Hypervisor/Node resource view: name=ubuntu-focal-rax-dfw-0030919238 | |
| 14:10:18 | sean-k-mooney | nope they all seem to line up | |
| 14:10:43 | sean-k-mooney | odd | |
| 14:11:07 | dansmith | do we do a disabled or service liveness check before we let you unshelve there? maybe the compute is stuck and not updating its counter? | |
| 14:12:15 | gibi | when we query the service we check for that it is up and enabled | |
| 14:12:21 | gmann | dansmith: you mean just before unshelve or while selecting the host? while selecting host we do check service is up and enable | |
| 14:12:43 | gmann | but after that shelve happen and then unshelve and that time no check before unshelve | |
| 14:13:02 | dansmith | gmann: yeah I mean whatever check we're doing that gives that error | |
| 14:13:15 | dansmith | is there a tempest test that is disabling a compute by chance for testing? | |
| 14:13:18 | sean-k-mooney | its proably worth checkign the schduler logs | |
| 14:13:33 | dansmith | oh is this just the scheduler "is this host okay" check? | |
| 14:15:20 | gmann | I do not think we have any such test of touching the compute services enable/disable | |
| 14:15:30 | gibi | it is not about sevice state | |
| 14:15:39 | gibi | the unshelve fails on compute_node_get_all_by_host | |
| 14:15:48 | gibi | so the compute is not in the DBV | |
| 14:15:49 | gibi | DB | |
| 14:17:15 | gibi | https://github.com/openstack/nova/blob/master/nova/compute/api.py#L4577 | |
| 14:17:21 | gibi | this where the unshelve fails | |
| 14:18:53 | sean-k-mooney | so it is acutlly trying to shelve and unshlve multiple times | |
| 14:19:15 | sean-k-mooney | it shleve offloaded and shelve to the compute | |
| 14:19:18 | sean-k-mooney | then it shelved again | |
| 14:19:28 | sean-k-mooney | and failed to unshleve to the contoler with the not found issue | |
| 14:19:52 | gibi | yes | |
| 14:20:12 | gibi | first it unselves back to where it was, then it unshelves to the other compute | |
| 14:20:42 | sean-k-mooney | and req-b3e7705d-a20b-42b3-8eb5-594181a096c8 is the request that fialed | |
| 14:20:52 | sean-k-mooney | which failed in the api right | |
| 14:20:57 | sean-k-mooney | its not in the secheuler | |
| 14:21:02 | gibi | yes | |
| 14:21:07 | gmann | yes it is failing in 2nd time unshelve on another host | |
| 14:21:10 | gibi | it is not fails on the scheduler | |
| 14:21:16 | gibi | it fails in the compute api | |
| 14:21:17 | gibi | https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/controller/logs/screen-n-api.txt | |
| 14:21:20 | gibi | https://github.com/openstack/nova/blob/master/nova/compute/api.py#L4577 | |
| 14:21:33 | gibi | unshelve checks is the requested host exists | |
| 14:21:49 | gibi | before it goes to the conductor | |
| 14:21:56 | sean-k-mooney | by calling compute_node_get_all_by_host | |
| 14:22:20 | gibi | it calls get_first_node_by_host_for_old_compat | |
| 14:22:26 | sean-k-mooney | thats on the main db i.e. cell db right | |
| 14:22:39 | dansmith | so that is looking up by the service, and getting the first compute node on the service | |
| 14:22:45 | dansmith | not by_host_and_node | |
| 14:23:13 | sean-k-mooney | for libvirt there should only be one but ya | |
| 14:23:32 | sean-k-mooney | if shelve was a thing for vmware or ironic that would not be ideal | |
| 14:23:38 | dansmith | sean-k-mooney: right but this is very loose association which doesn't have FK constraints | |
| 14:24:05 | sean-k-mooney | why is it hitting the cell db. it could look at the host mapping table in the api db right | |
| 14:24:14 | sean-k-mooney | or cell mappings table | |
| 14:24:21 | dansmith | eh? | |
| 14:24:29 | dansmith | those just tell us which db the host is in | |
| 14:24:43 | sean-k-mooney | right but if all you care about is does it exist is that not enough | |
| 14:24:51 | gibi | sean-k-mooney: you have a point, so the compute-api does a cell DB query without targeting the cell first | |
| 14:24:51 | dansmith | nothing about whether it's really there or up or whatever | |
| 14:25:23 | sean-k-mooney | dansmith: i guess it could be down but im not sure this is checkign that | |
| 14:25:28 | dansmith | sean-k-mooney: or deleted | |
| 14:25:48 | sean-k-mooney | i would expct that to be check by the schduler to be honest | |
| 14:25:53 | dansmith | gibi: these hosts are in the same cell though so if that wasn't working, we wouldn't have found the first one right? | |
| 14:26:22 | gibi | this is the multicell job so I assume we have hosts in different cells | |
| 14:26:45 | sean-k-mooney | yes there are only two nodes in the job | |
| 14:26:51 | sean-k-mooney | so each is in a different cell | |
| 14:26:59 | dansmith | multicell meaning superconductor not actually one host per cell right? | |
| 14:27:27 | sean-k-mooney | well we can check but i tought this was 2 cells | |
| 14:27:28 | dansmith | or is this really unshelving to different cells? | |
| 14:27:50 | dansmith | can we unshelve across cells? | |
| 14:27:52 | dansmith | I thought not | |
| 14:28:16 | dansmith | so maybe we're targeted at the cell where the instance is and not the cell where the host is that we're unshelving to? | |
| 14:28:24 | dansmith | not sure how that would have ever passed | |
| 14:28:48 | sean-k-mooney | the compute is not runnign a second conductor | |
| 14:28:55 | gibi | I'm looking at the compute.api.API.unshelve and I think the context is not targeted | |
| 14:29:14 | sean-k-mooney | oh the contoler is ruuning both cell1 and cell2 conductors | |
| 14:29:36 | sean-k-mooney | and the supper conductor | |
| 14:30:24 | sean-k-mooney | https://de836787b7e59a5adc13-298f4365cc798f0001a632f171eb41d6.ssl.cf2.rackcdn.com/831219/22/check/nova-multi-cell/9d8aa66/compute1/logs/etc/nova/nova_cell1_conf.txt | |
| 14:30:51 | sean-k-mooney | transport_url = rabbit://stackrabbit:secretrabbit@10.208.192.130:5672/nova_cell2 | |
| 14:31:00 | sean-k-mooney | so ya compute1 is nova_cell2 | |
| 14:31:20 | sean-k-mooney | and contoler is transport_url = rabbit://stackrabbit:secretrabbit@10.208.192.130:5672/nova_cell1 | |