| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-05 | |||
| 19:32:43 | melwitt | so if someone messes up and deletes a service and then says oops that was a mistake, then all those instances are expected not to work? | |
| 19:32:45 | sean-k-mooney | its a uuid4 and not based on the hostname/hypervior_hostname | |
| 19:32:52 | melwitt | dang | |
| 19:32:53 | dansmith | you can't delete a service with instances on it | |
| 19:33:10 | melwitt | ok, so that saves it I guess? ok | |
| 19:33:23 | sean-k-mooney | dansmith: are you sure | |
| 19:33:23 | dansmith | saves it from the single-click-mega-fail, but.. :) | |
| 19:33:27 | dansmith | pretty sure | |
| 19:33:52 | sean-k-mooney | ok cause i know we have code to loop over the allocation in placment and delete them before we delete the placment rp when teh compute serivce is deleted | |
| 19:34:05 | melwitt | just seems so harsh lol (if it were possible to delete the service while instances are on it) | |
| 19:34:07 | dansmith | yup | |
| 19:34:20 | sean-k-mooney | i guess that is just to prevent leaked allocation blocking the placment cleanup | |
| 19:35:37 | sean-k-mooney | ah https://github.com/openstack/nova/blob/master/nova/api/openstack/compute/services.py#L269-L282 | |
| 19:35:43 | sean-k-mooney | we special case the nova-compute | |
| 19:36:02 | sean-k-mooney | so ya you cant delete it if it has instance | |
| 19:36:18 | melwitt | ok, well, if that's the case then I understand why and agree the test can be removed entirely. just seems so harsh, if what I was thinking were possible (and it is not possible bc we don't let you delete a service with instances mapped to it) | |
| 19:36:19 | sean-k-mooney | in which case provide the placment clean up happens properly it does not really matter if the uuid changes in that case | |
| 19:36:28 | sean-k-mooney | or if we undelete | |
| 19:37:07 | sean-k-mooney | https://github.com/openstack/nova/commit/42f62f1ed2ad76829eb9d40a8b9646a523f6381f | |
| 19:37:25 | sean-k-mooney | melwitt: it was only blokced in rocky it looks like | |
| 19:38:03 | sean-k-mooney | https://bugs.launchpad.net/nova/+bug/1763183 | |
| 19:38:13 | melwitt | I think we (maybe I) backported it downstream | |
| 19:38:36 | sean-k-mooney | well it was backported upstream to pike | |
| 19:38:39 | melwitt | I just was not thinking about it or remembering it | |
| 19:38:47 | melwitt | ah ok | |
| 19:39:56 | sean-k-mooney | i rememebr being able to delete compute serivce with instance at one point but i feel like that is just because i mess up my local devstack not because i planed to do it | |
| 19:40:08 | dansmith | melwitt: here are the most service-delete-y tests we have in functional/ https://github.com/openstack/nova/blob/master/nova/tests/functional/wsgi/test_services.py#L119 | |
| 19:40:11 | melwitt | yeah you used to be able to | |
| 19:40:18 | dansmith | none of them ensure we can start an instance on the resurrected service, | |
| 19:40:28 | dansmith | although they do restart the compute to make sure it comes back up | |
| 19:40:43 | dansmith | which is the thing sean-k-mooney and laugh at outside a fake environment :P | |
| 19:41:08 | melwitt | I see, ok. thanks | |
| 19:41:22 | dansmith | melwitt: so your demand is me adding a test that a resurrected compute can fake boot a fake instance un a fake environment, and then I can delete this regression test, right? | |
| 19:41:26 | dansmith | (snarky on purpose, but serious) | |
| 19:41:45 | melwitt | sorry for the longer convo. I was very confused by the test and then I was erroneously thinking of an accidental service delete scenario | |
| 19:42:04 | dansmith | don't apologize | |
| 19:42:20 | melwitt | yeah, I said earlier I understand now and agree the test can be removed without loss of anything | |
| 19:42:24 | dansmith | the stuff I'm having to do in this set to make such a simple thing work is ridiculously incestuous | |
| 19:42:59 | dansmith | melwitt: well, I think adding a "and can boot something" thing to those ^ would make that a defensible position for me :) | |
| 19:43:01 | melwitt | I bet :\ | |
| 19:44:12 | melwitt | thanks for that 😂 | |
| 19:44:12 | melwitt | thanks for that 😂 | |
| 19:49:54 | sean-k-mooney | dansmith: alot of that likel come form how the fixture make restarting compute service work in the past | |
| 19:50:07 | dansmith | yes, I'm well aware | |
| 19:50:33 | melwitt | dansmith: I agree adding a "and can boot something" to those existing tests is a nice thing to cover. but I don't expect it to have to be part of your series, to be clear | |
| 19:51:02 | sean-k-mooney | with the stable uuid serise i am assuming you will have a functional test that start with an empty db and starts a comptue service with the uuid specifed in a file | |
| 19:51:33 | sean-k-mooney | you have a seperte test that delete it form teh db and starts it again if you wanted | |
| 19:52:51 | sean-k-mooney | but ya i think we agreed on nuke the thing and move on with your seriese | |
| 19:53:36 | melwitt | yes | |
| 19:55:05 | dansmith | well, I figure I need to add the other when I drop the regression test | |
| 19:55:21 | dansmith | there's something weird though about not seeing the provider get recreated after restarting the old compute, | |
| 19:55:27 | dansmith | although I see it happen in the logs | |
| 19:56:37 | sean-k-mooney | that happens after teh perodic task runs although it also happens i think in init host | |
| 19:57:06 | dansmith | I see it created before I look for it | |
| 19:57:29 | dansmith | https://pastebin.com/isqJXnfW | |
| 19:57:43 | dansmith | first line is it being created in our db, then placement, then the last one is looking for it, but it's missing | |
| 19:59:08 | dansmith | I kinda wonder if there's a bug causing us to find the old deleted compute node before the new one, and then return nothing because it's deleted | |
| 20:02:09 | dansmith | hah | |
| 20:02:10 | dansmith | 2023-01-05 12:01:58,067 INFO [nova.api.openstack.compute.hypervisors] Unable to find service for compute node host1. The service may be deleted and compute nodes need to be manually cleaned up. | |
| 20:02:37 | dansmith | that's what happens when I try to list hypervisors with the old name after re-starting the service | |
| 20:02:48 | dansmith | the service object should be undeleted, a new compute node was created, | |
| 20:03:01 | dansmith | yet listing doesn't include *either* because of that ^ | |
| 20:03:18 | dansmith | melwitt: see what we mean now? :) | |
| 20:03:56 | melwitt | 😵💫 | |
| 20:04:45 | sean-k-mooney | could this be related to the cell mappings | |
| 20:04:51 | sean-k-mooney | in the api db | |
| 20:05:25 | sean-k-mooney | as in does discover host need to be run | |
| 20:06:10 | dansmith | god I hope not | |
| 20:06:24 | sean-k-mooney | dansmith: by the way i do know that if the resouce tracker is broken the compute service can show up in the comptue service list but the compute node will not show up in the hypervior list | |
| 20:06:48 | sean-k-mooney | so if you run the test with OS_DEBUG maybe there is somethign breaking in the restart | |
| 20:07:36 | sean-k-mooney | i only see info logs in the output you pasted so fi this is from a functional test then you might need OS_DEBUG=1 | |
| 20:08:14 | dansmith | sure enough: Host 'host1' is not mapped to any cell | |
| 20:08:17 | sean-k-mooney | although if it was broken that way i woudl expect to see some trace backs or Error logs so debug should not be required | |
| 20:09:53 | dansmith | OS_DEBUG changed lately btw | |
| 20:10:03 | dansmith | I used to set OS_DEBUG=y but that doesn't work anymore | |
| 20:10:12 | dansmith | is =1 the new magic? | |
| 20:11:06 | sean-k-mooney | i have always used 1 but not sure if/when that changed | |
| 20:11:19 | sean-k-mooney | i dont think its every really been documented properly | |
| 20:14:46 | dansmith | yeah | |
| 20:14:51 | dansmith | =y generates an exception now | |
| 20:16:45 | sean-k-mooney | i assuem its anythign loosely equivalent to true in a c like language | |
| 20:17:52 | sean-k-mooney | for the cell mappings stuff i dont think we normally run discovier hosts explictly anywhere in our funct tests | |
| 20:18:13 | dansmith | I didn't either, which is why it seems weird to me that it fails like that | |
| 20:18:22 | dansmith | maybe we insert the mapping in start but not in restart? | |
| 20:18:27 | dansmith | anyway, | |
| 20:18:32 | dansmith | I'll leave that as s #FIXME for later | |
| 20:18:33 | sean-k-mooney | it might be burried in some of the compute create code but ill admit i have neverlooked | |
| 20:18:57 | sean-k-mooney | you could always cheat with the conductor periodic if you needed to in the short term | |
| 20:19:05 | sean-k-mooney | anywya im going to call it a day soon | |
| 20:19:22 | dansmith | if I don't verify the new rp I'll make it past | |
| 20:20:05 | sean-k-mooney | i dont see how the cell mappings stuff could impact the palcment part by the way. what was the exception you got? | |
| 20:20:41 | sean-k-mooney | the cell mappiing shoudl only affect calling the comptue service via rpc | |
| 20:21:42 | sean-k-mooney | so the rp thing most be somethign else | |
| 20:53:13 | dansmith | it impacts the placement stuff only in the verification in the tests, because we use hypervisors to find the rp uuid and then check the allocations | |
| 20:53:23 | dansmith | if I just don't do that validation (like other parts of the test) them I'm good | |
| #openstack-nova - 2023-01-06 | |||
| 02:35:16 | opendevreview | Nobuhiro MIKI proposed openstack/nova-specs master: Add PXB support for libvirt https://review.opendev.org/c/openstack/nova-specs/+/869416 | |
| 11:03:28 | opendevreview | Aaron S proposed openstack/nova master: Add further workaround features for qemu_monitor_announce_self https://review.opendev.org/c/openstack/nova/+/867324 | |
| 14:46:26 | stephenfin | gibi: No point rechecking jobs that exhibit this failure | |
| 14:46:34 | stephenfin | tox.tox_env.python.api.NoInterpreter: could not find python interpreter matching any of the specs functional-py39 | |
| 14:46:36 | gibi | ahh | |