| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-05 | |||
| 19:38:13 | melwitt | I think we (maybe I) backported it downstream | |
| 19:38:36 | sean-k-mooney | well it was backported upstream to pike | |
| 19:38:39 | melwitt | I just was not thinking about it or remembering it | |
| 19:38:47 | melwitt | ah ok | |
| 19:39:56 | sean-k-mooney | i rememebr being able to delete compute serivce with instance at one point but i feel like that is just because i mess up my local devstack not because i planed to do it | |
| 19:40:08 | dansmith | melwitt: here are the most service-delete-y tests we have in functional/ https://github.com/openstack/nova/blob/master/nova/tests/functional/wsgi/test_services.py#L119 | |
| 19:40:11 | melwitt | yeah you used to be able to | |
| 19:40:18 | dansmith | none of them ensure we can start an instance on the resurrected service, | |
| 19:40:28 | dansmith | although they do restart the compute to make sure it comes back up | |
| 19:40:43 | dansmith | which is the thing sean-k-mooney and laugh at outside a fake environment :P | |
| 19:41:08 | melwitt | I see, ok. thanks | |
| 19:41:22 | dansmith | melwitt: so your demand is me adding a test that a resurrected compute can fake boot a fake instance un a fake environment, and then I can delete this regression test, right? | |
| 19:41:26 | dansmith | (snarky on purpose, but serious) | |
| 19:41:45 | melwitt | sorry for the longer convo. I was very confused by the test and then I was erroneously thinking of an accidental service delete scenario | |
| 19:42:04 | dansmith | don't apologize | |
| 19:42:20 | melwitt | yeah, I said earlier I understand now and agree the test can be removed without loss of anything | |
| 19:42:24 | dansmith | the stuff I'm having to do in this set to make such a simple thing work is ridiculously incestuous | |
| 19:42:59 | dansmith | melwitt: well, I think adding a "and can boot something" thing to those ^ would make that a defensible position for me :) | |
| 19:43:01 | melwitt | I bet :\ | |
| 19:44:12 | melwitt | thanks for that 😂 | |
| 19:44:12 | melwitt | thanks for that 😂 | |
| 19:49:54 | sean-k-mooney | dansmith: alot of that likel come form how the fixture make restarting compute service work in the past | |
| 19:50:07 | dansmith | yes, I'm well aware | |
| 19:50:33 | melwitt | dansmith: I agree adding a "and can boot something" to those existing tests is a nice thing to cover. but I don't expect it to have to be part of your series, to be clear | |
| 19:51:02 | sean-k-mooney | with the stable uuid serise i am assuming you will have a functional test that start with an empty db and starts a comptue service with the uuid specifed in a file | |
| 19:51:33 | sean-k-mooney | you have a seperte test that delete it form teh db and starts it again if you wanted | |
| 19:52:51 | sean-k-mooney | but ya i think we agreed on nuke the thing and move on with your seriese | |
| 19:53:36 | melwitt | yes | |
| 19:55:05 | dansmith | well, I figure I need to add the other when I drop the regression test | |
| 19:55:21 | dansmith | there's something weird though about not seeing the provider get recreated after restarting the old compute, | |
| 19:55:27 | dansmith | although I see it happen in the logs | |
| 19:56:37 | sean-k-mooney | that happens after teh perodic task runs although it also happens i think in init host | |
| 19:57:06 | dansmith | I see it created before I look for it | |
| 19:57:29 | dansmith | https://pastebin.com/isqJXnfW | |
| 19:57:43 | dansmith | first line is it being created in our db, then placement, then the last one is looking for it, but it's missing | |
| 19:59:08 | dansmith | I kinda wonder if there's a bug causing us to find the old deleted compute node before the new one, and then return nothing because it's deleted | |
| 20:02:09 | dansmith | hah | |
| 20:02:10 | dansmith | 2023-01-05 12:01:58,067 INFO [nova.api.openstack.compute.hypervisors] Unable to find service for compute node host1. The service may be deleted and compute nodes need to be manually cleaned up. | |
| 20:02:37 | dansmith | that's what happens when I try to list hypervisors with the old name after re-starting the service | |
| 20:02:48 | dansmith | the service object should be undeleted, a new compute node was created, | |
| 20:03:01 | dansmith | yet listing doesn't include *either* because of that ^ | |
| 20:03:18 | dansmith | melwitt: see what we mean now? :) | |
| 20:03:56 | melwitt | 😵💫 | |
| 20:04:45 | sean-k-mooney | could this be related to the cell mappings | |
| 20:04:51 | sean-k-mooney | in the api db | |
| 20:05:25 | sean-k-mooney | as in does discover host need to be run | |
| 20:06:10 | dansmith | god I hope not | |
| 20:06:24 | sean-k-mooney | dansmith: by the way i do know that if the resouce tracker is broken the compute service can show up in the comptue service list but the compute node will not show up in the hypervior list | |
| 20:06:48 | sean-k-mooney | so if you run the test with OS_DEBUG maybe there is somethign breaking in the restart | |
| 20:07:36 | sean-k-mooney | i only see info logs in the output you pasted so fi this is from a functional test then you might need OS_DEBUG=1 | |
| 20:08:14 | dansmith | sure enough: Host 'host1' is not mapped to any cell | |
| 20:08:17 | sean-k-mooney | although if it was broken that way i woudl expect to see some trace backs or Error logs so debug should not be required | |
| 20:09:53 | dansmith | OS_DEBUG changed lately btw | |
| 20:10:03 | dansmith | I used to set OS_DEBUG=y but that doesn't work anymore | |
| 20:10:12 | dansmith | is =1 the new magic? | |
| 20:11:06 | sean-k-mooney | i have always used 1 but not sure if/when that changed | |
| 20:11:19 | sean-k-mooney | i dont think its every really been documented properly | |
| 20:14:46 | dansmith | yeah | |
| 20:14:51 | dansmith | =y generates an exception now | |
| 20:16:45 | sean-k-mooney | i assuem its anythign loosely equivalent to true in a c like language | |
| 20:17:52 | sean-k-mooney | for the cell mappings stuff i dont think we normally run discovier hosts explictly anywhere in our funct tests | |
| 20:18:13 | dansmith | I didn't either, which is why it seems weird to me that it fails like that | |
| 20:18:22 | dansmith | maybe we insert the mapping in start but not in restart? | |
| 20:18:27 | dansmith | anyway, | |
| 20:18:32 | dansmith | I'll leave that as s #FIXME for later | |
| 20:18:33 | sean-k-mooney | it might be burried in some of the compute create code but ill admit i have neverlooked | |
| 20:18:57 | sean-k-mooney | you could always cheat with the conductor periodic if you needed to in the short term | |
| 20:19:05 | sean-k-mooney | anywya im going to call it a day soon | |
| 20:19:22 | dansmith | if I don't verify the new rp I'll make it past | |
| 20:20:05 | sean-k-mooney | i dont see how the cell mappings stuff could impact the palcment part by the way. what was the exception you got? | |
| 20:20:41 | sean-k-mooney | the cell mappiing shoudl only affect calling the comptue service via rpc | |
| 20:21:42 | sean-k-mooney | so the rp thing most be somethign else | |
| 20:53:13 | dansmith | it impacts the placement stuff only in the verification in the tests, because we use hypervisors to find the rp uuid and then check the allocations | |
| 20:53:23 | dansmith | if I just don't do that validation (like other parts of the test) them I'm good | |
| #openstack-nova - 2023-01-06 | |||
| 02:35:16 | opendevreview | Nobuhiro MIKI proposed openstack/nova-specs master: Add PXB support for libvirt https://review.opendev.org/c/openstack/nova-specs/+/869416 | |
| 11:03:28 | opendevreview | Aaron S proposed openstack/nova master: Add further workaround features for qemu_monitor_announce_self https://review.opendev.org/c/openstack/nova/+/867324 | |
| 14:46:26 | stephenfin | gibi: No point rechecking jobs that exhibit this failure | |
| 14:46:34 | stephenfin | tox.tox_env.python.api.NoInterpreter: could not find python interpreter matching any of the specs functional-py39 | |
| 14:46:36 | gibi | ahh | |
| 14:46:40 | stephenfin | it's another tox 4 bug | |
| 14:46:44 | gibi | nice | |
| 14:46:50 | stephenfin | https://github.com/tox-dev/tox/issues/2811 | |
| 14:47:31 | gibi | what can we do? | |
| 14:47:32 | stephenfin | I've happened to expose it by fixing another bug that resulted in us using the wrong interpreter version | |
| 14:47:38 | stephenfin | :( | |
| 14:48:06 | sean-k-mooney | i have been seeing it since before you fix | |
| 14:48:23 | sean-k-mooney | but ya all the gates are currently blocked | |
| 14:48:29 | stephenfin | yeah, most likely on projects without base_python set | |
| 14:49:13 | sean-k-mooney | i saw it on nova yesterday and i think on older builds form durign the week | |
| 14:49:18 | sean-k-mooney | we have base_python set | |
| 14:49:43 | sean-k-mooney | we dont actully need to have it set anymore since we are python3 only | |
| 14:49:47 | stephenfin | the fix to tox merged yesterday so it was probably that | |
| 14:50:19 | sean-k-mooney | ya the release happend 19 hours ago but i toughthe builds were older then that | |
| 14:50:26 | sean-k-mooney | i saw it on gibis seriese | |
| 14:51:23 | sean-k-mooney | im wondering if we shoudl repin to tox <4.0 tempoerally | |
| 14:52:02 | sean-k-mooney | we can proably wait another week but if we cant resolve the issue by the end of next week i think we should | |
| 14:59:10 | gibi | sean-k-mooney, stephenfin: is there a mail thread about the gate block on the ML yet or should I send one? | |
| 14:59:35 | stephenfin | There isn't. The fix is here https://github.com/tox-dev/tox/pull/2828 though if you want to send one and point to that | |
| 14:59:59 | gibi | I will send one | |
| 15:01:35 | dansmith | sean-k-mooney: tox has been slowly breaking everything for weeks now.. pinning to <4 temporarily seems futile | |