| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-27 | |||
| 18:40:31 | dansmith | I really think that if you take out the "randomly deleted a key state file from the system" then it prevents quite a bit | |
| 18:41:07 | dansmith | how is this a new db exception? it's the same db exception as before if we try to create a conflicting compute node record | |
| 18:41:11 | sean-k-mooney | dansmith: well if your relase automation updated the uuid it would cause the same failure im seeing | |
| 18:41:14 | sean-k-mooney | that was basiclaly step 4 | |
| 18:41:30 | sean-k-mooney | no | |
| 18:41:36 | sean-k-mooney | before we looked it up by hostname | |
| 18:41:43 | sean-k-mooney | and would not have got a colliion in this case | |
| 18:42:00 | sean-k-mooney | so we are trying to create a duplicte record that would not have been created before | |
| 18:42:05 | dansmith | okay I'm getting frustrated | |
| 18:42:20 | dansmith | shall we take this to a gmeet? | |
| 18:42:29 | sean-k-mooney | sure | |
| 18:42:39 | sean-k-mooney | im not trying to frusttrate you by the way | |
| 18:42:49 | sean-k-mooney | just letting you knwo what im finding | |
| 18:42:52 | dansmith | meet.google.com/gkf-fdhr-wgr | |
| 19:26:16 | dansmith | sean-k-mooney: one other thing, we could also assert that if we're already upgraded *and* are not starting fresh, we could abort if the uuid file is missing | |
| 19:26:22 | dansmith | i.e. your didn't-bind-mount case | |
| 19:26:45 | dansmith | although, | |
| 19:27:08 | dansmith | if we handle that in the extra check we discussed, we can say "found X expected Y" which will be the easy way for them to fix their stuff, even if X is empty | |
| 19:27:52 | sean-k-mooney | yes i thought that was one of the things i said above. | |
| 19:28:10 | sean-k-mooney | yep | |
| 19:28:22 | dansmith | oh, maybe I was foaming at the mouth and missed it | |
| 19:30:35 | sean-k-mooney | i read over your comments on teh persist change too so im ok to proceed with that now based on what we discussed | |
| 19:31:02 | dansmith | ack | |
| 19:31:10 | sean-k-mooney | so the first 4 have +w and the first 2 are merged | |
| 19:32:45 | dansmith | thanks, I'll wait until those merge or fail before I push anything else up | |
| 19:33:39 | sean-k-mooney | ok im going to see what else i can break or not break and ill review the last 3 on monday | |
| 19:53:53 | opendevreview | Merged openstack/nova master: Add get_available_node_uuids() to virt driver https://review.opendev.org/c/openstack/nova/+/863917 | |
| 20:04:59 | sean-k-mooney | dansmith: fyi im gong to bold the titles of the tests that look odd | |
| 20:05:58 | sean-k-mooney | but the rename logic does not detact a change in hypervior_hostname today as long as CONF.host does not change | |
| 20:06:07 | sean-k-mooney | https://etherpad.opendev.org/p/Stable-compute-uuid-manual-testing#L216 | |
| 20:06:43 | dansmith | and that's because we get that from libvirt yeah? | |
| 20:06:54 | sean-k-mooney | yep | |
| 20:07:27 | sean-k-mooney | so if the value of virsh hostname changes then we update hypervior_hostname in the db and rename the placemnet RP | |
| 20:07:57 | dansmith | and that's problematic why, just because cinder/neutron will be unhappy? | |
| 20:08:21 | sean-k-mooney | im sure that will break things like QOS or other nested resouce providers ya or at least how we discover them | |
| 20:08:30 | dansmith | that's technically out of scope of what I was trying to guard against in this set | |
| 20:08:35 | sean-k-mooney | that siad if they also rename there RPs then it might just fix itslef | |
| 20:08:35 | dansmith | and what if you change it back? | |
| 20:08:53 | sean-k-mooney | thats going to be the next test | |
| 20:08:57 | dansmith | okay | |
| 20:09:11 | sean-k-mooney | after that im going to try changing the value via /etc/hosts | |
| 20:09:21 | sean-k-mooney | i expect changing it back to revert all the chagnes | |
| 20:10:00 | sean-k-mooney | i proably shoudl test this with nested resouce provides form neutron but ill do that after all the simiple tests | |
| 20:11:07 | sean-k-mooney | actully i need to check a few db tables first | |
| 20:13:00 | dansmith | okay, so you're saying it updates plaacement but does *not* update our compute node right? | |
| 20:13:24 | sean-k-mooney | it updates teh cn table and placment they are both kept in sync | |
| 20:13:33 | sean-k-mooney | im checking what the api db looks like | |
| 20:14:51 | dansmith | L219 says otherwise | |
| 20:15:00 | dansmith | or you meant "isn't detected to abort" | |
| 20:15:56 | sean-k-mooney | i think we shoudl abort starting up if the hypervior_hostname for a compute node changes yes | |
| 20:16:03 | sean-k-mooney | at least for libvirt | |
| 20:16:30 | sean-k-mooney | but again thats an extra check we can add to the end fo the series | |
| 20:17:15 | sean-k-mooney | so i proably should have a vm runnign while i do this as i think the instance.node is going to be out of sync | |
| 20:17:22 | sean-k-mooney | so ill boot one before i fix it | |
| 20:17:58 | sean-k-mooney | and my expectation is that when i restor the old hostname the isntance.node will not be updated but placement and the compute nodes table will be | |
| 20:18:10 | sean-k-mooney | thats proably the last test ill do tonight | |
| 20:18:18 | sean-k-mooney | but ill do more on monday. | |
| 20:18:25 | dansmith | yeah, like I said, out of scope of what I'm trying to do, but probably should be in scope | |
| 20:18:27 | dansmith | ack | |
| 20:19:23 | sean-k-mooney | actully i need to test with provider.yaml too at some point | |
| 20:19:41 | sean-k-mooney | we can refernce the RP by name i belive which uses the hypervior_hostname | |
| 20:19:48 | sean-k-mooney | we can also refrrence it by UUID | |
| 20:20:05 | sean-k-mooney | so i suspect the uuid way will be fine but the other way might break | |
| 20:20:20 | sean-k-mooney | again not nessiarly in scope but i want to know what happens | |
| 20:41:27 | sean-k-mooney | ok so if there are allocation in placment we cant delete and rename the resouce provider becaues we get a 409 | |
| 20:42:08 | sean-k-mooney | so that fails and we never update the compute node record either | |
| 20:42:32 | sean-k-mooney | the perodic is broken | |
| 20:43:24 | sean-k-mooney | but the agent does not exit and we just get tracebacks in the logs. | |
| 20:44:12 | sean-k-mooney | ill leave it there for today and come back to this on monday | |
| 20:46:34 | dansmith | wait, | |
| 20:46:43 | dansmith | I thought you said it *did* update the placement provider hostname? | |
| 20:47:56 | sean-k-mooney | it appeard too but based on that log it deleted the RP and recreated it the the old uuid and new name | |
| 20:48:17 | sean-k-mooney | so its not updating in place its trying too delete orphan compute and then creating a new one | |
| 20:48:27 | sean-k-mooney | that by the way i think is ironic code | |
| 20:48:52 | sean-k-mooney | or rather code we have in the common manager loop for ironic | |
| 20:50:39 | sean-k-mooney | we we detech hypervior_hostname changes and abort just like conf.host we dont have to care about this | |
| 20:50:53 | dansmith | ah because it had no instances? | |
| 20:51:04 | sean-k-mooney | right so with no instnaces it deelted it fine | |
| 20:51:09 | dansmith | right okay | |
| 20:51:18 | sean-k-mooney | the 409 conflict in placment is because fo the allcoation for the instnace | |
| 20:51:23 | dansmith | so if no instances, maybe no harm to cinder and neutron? | |
| 20:51:40 | sean-k-mooney | it might still break the naming | |
| 20:51:51 | sean-k-mooney | but it might be ok in the no instance case | |
| 20:52:12 | sean-k-mooney | i will confirure bandwith QOS or something next week and see | |
| 20:52:23 | sean-k-mooney | cindier i dont think will use placment at all reight now | |
| 20:52:27 | sean-k-mooney | but cyborg could break | |
| 20:53:15 | sean-k-mooney | cyborg and neutron might need me to restarck so ill test what i can before that | |
| 20:54:02 | sean-k-mooney | dansmith: in this particalr case this placement exception i think happend before your code | |
| 20:54:20 | dansmith | during reshape or something? | |
| 20:54:24 | sean-k-mooney | i have defintly seen this before and its what i was expecting if your code did not block it | |
| 20:54:58 | dansmith | fml | |
| 20:54:58 | dansmith | so that sanity check of the hostname->node mapping generates 94 functional test failures | |
| 20:55:07 | sean-k-mooney | no i have seen this when people actully change dns/hostname but had CONF.host set | |
| 20:57:18 | sean-k-mooney | we dont need to fix all theses cases this cycle either like i know that test case 10 is a prexisitng failure mode | |
| 20:57:38 | sean-k-mooney | anyway im going to go get food | |
| 20:57:47 | sean-k-mooney | dont spend your weekend on this o/ | |
| 20:59:21 | dansmith | I shan't, you either | |
| #openstack-nova - 2023-01-28 | |||
| 19:58:06 | opendevreview | Takashi Natsume proposed openstack/nova-specs master: Create specs directory for 2023.2 Bobcat https://review.opendev.org/c/openstack/nova-specs/+/872068 | |
| #openstack-nova - 2023-01-30 | |||
| 09:27:45 | plibeau | hello guys, if you have sometime to review please: https://review.opendev.org/c/openstack/nova/+/861172 | |
| 11:17:01 | opendevreview | Merged openstack/nova master: Fix rescue volume-based instance https://review.opendev.org/c/openstack/nova/+/852737 | |