Earlier  
Posted Nick Remark
#openstack-nova - 2023-01-27
18:31:57 sean-k-mooney so i dont think we shoudl be ignoring the host/hypervior_hostname entirely espcially when we get the db duplicte key error
18:32:28 dansmith isn't that the entire point of this series? to break that as the link?
18:32:40 sean-k-mooney no
18:32:51 sean-k-mooney its to make sure a host never change its hostname or host value
18:32:56 sean-k-mooney that is not entirly the same thing
18:33:14 dansmith I totally disagree that that's the point of this series :)
18:33:14 sean-k-mooney we want to make sure we always have a stable way to mapp a host to the same comptue node record
18:33:29 sean-k-mooney and that the host and hypervior_hostname dont change
18:33:35 dansmith because I didn't even have the rename protection in there until late
18:34:18 sean-k-mooney well this would not have prevented the issues our downstream custoemr had without the host/hypersior host checks
18:34:58 dansmith it would, if they didn't delete the state file
18:35:25 dansmith like I said, if the uuid is actually a compute node in the db, but with a hostname that doesn't match, it will abort
18:35:36 sean-k-mooney and you want to rely on them not doing something we told them not to do when they are doing something else we told them not to do :)
18:35:57 dansmith all I want to do is make finding the compute node not tied to the hostname
18:36:18 sean-k-mooney ok so that will just result in it failing with placment
18:36:43 sean-k-mooney because we will move the dduplicate key to ther resouce provider crations
18:36:55 dansmith but we won't be recreating providers
18:37:06 dansmith we'll be accessing them by uuid, which hasn't changed right?
18:37:24 dansmith the hostname may be wrong, and if something else claims the hostname, then it will fail to create one,
18:37:31 sean-k-mooney well im about to start testing that stuff so we will see
18:37:43 dansmith but the case we've had downstream was not that they *shuffled* their compute node names, but rather moved to a different naming schema, still no overlaps
18:37:52 sean-k-mooney setting it back to the correct uuid and ill change the CONF.HOST and hostname seperatly
18:38:46 sean-k-mooney anyway my inital feedback is this is not preventign as much as i was expecting it to
18:38:47 dansmith so here's the accepted spec: https://specs.openstack.org/openstack/nova-specs/specs/2023.1/approved/stable-compute-uuid.html#proposed-change
18:38:57 dansmith and I think that's covered here
18:39:09 dansmith it says that the uuid file will be what we use to find the compute node,
18:39:26 dansmith and we will detect compute node renames
18:40:29 sean-k-mooney right and i kind fo assume "we will not intoduce any new DB exctpions that prevent the resouce tracker form working" woudl be an imporant point too
18:40:31 dansmith I really think that if you take out the "randomly deleted a key state file from the system" then it prevents quite a bit
18:41:07 dansmith how is this a new db exception? it's the same db exception as before if we try to create a conflicting compute node record
18:41:11 sean-k-mooney dansmith: well if your relase automation updated the uuid it would cause the same failure im seeing
18:41:14 sean-k-mooney that was basiclaly step 4
18:41:30 sean-k-mooney no
18:41:36 sean-k-mooney before we looked it up by hostname
18:41:43 sean-k-mooney and would not have got a colliion in this case
18:42:00 sean-k-mooney so we are trying to create a duplicte record that would not have been created before
18:42:05 dansmith okay I'm getting frustrated
18:42:20 dansmith shall we take this to a gmeet?
18:42:29 sean-k-mooney sure
18:42:39 sean-k-mooney im not trying to frusttrate you by the way
18:42:49 sean-k-mooney just letting you knwo what im finding
18:42:52 dansmith meet.google.com/gkf-fdhr-wgr
19:26:16 dansmith sean-k-mooney: one other thing, we could also assert that if we're already upgraded *and* are not starting fresh, we could abort if the uuid file is missing
19:26:22 dansmith i.e. your didn't-bind-mount case
19:26:45 dansmith although,
19:27:08 dansmith if we handle that in the extra check we discussed, we can say "found X expected Y" which will be the easy way for them to fix their stuff, even if X is empty
19:27:52 sean-k-mooney yes i thought that was one of the things i said above.
19:28:10 sean-k-mooney yep
19:28:22 dansmith oh, maybe I was foaming at the mouth and missed it
19:30:35 sean-k-mooney i read over your comments on teh persist change too so im ok to proceed with that now based on what we discussed
19:31:02 dansmith ack
19:31:10 sean-k-mooney so the first 4 have +w and the first 2 are merged
19:32:45 dansmith thanks, I'll wait until those merge or fail before I push anything else up
19:33:39 sean-k-mooney ok im going to see what else i can break or not break and ill review the last 3 on monday
19:53:53 opendevreview Merged openstack/nova master: Add get_available_node_uuids() to virt driver https://review.opendev.org/c/openstack/nova/+/863917
20:04:59 sean-k-mooney dansmith: fyi im gong to bold the titles of the tests that look odd
20:05:58 sean-k-mooney but the rename logic does not detact a change in hypervior_hostname today as long as CONF.host does not change
20:06:07 sean-k-mooney https://etherpad.opendev.org/p/Stable-compute-uuid-manual-testing#L216
20:06:43 dansmith and that's because we get that from libvirt yeah?
20:06:54 sean-k-mooney yep
20:07:27 sean-k-mooney so if the value of virsh hostname changes then we update hypervior_hostname in the db and rename the placemnet RP
20:07:57 dansmith and that's problematic why, just because cinder/neutron will be unhappy?
20:08:21 sean-k-mooney im sure that will break things like QOS or other nested resouce providers ya or at least how we discover them
20:08:30 dansmith that's technically out of scope of what I was trying to guard against in this set
20:08:35 dansmith and what if you change it back?
20:08:35 sean-k-mooney that siad if they also rename there RPs then it might just fix itslef
20:08:53 sean-k-mooney thats going to be the next test
20:08:57 dansmith okay
20:09:11 sean-k-mooney after that im going to try changing the value via /etc/hosts
20:09:21 sean-k-mooney i expect changing it back to revert all the chagnes
20:10:00 sean-k-mooney i proably shoudl test this with nested resouce provides form neutron but ill do that after all the simiple tests
20:11:07 sean-k-mooney actully i need to check a few db tables first
20:13:00 dansmith okay, so you're saying it updates plaacement but does *not* update our compute node right?
20:13:24 sean-k-mooney it updates teh cn table and placment they are both kept in sync
20:13:33 sean-k-mooney im checking what the api db looks like
20:14:51 dansmith L219 says otherwise
20:15:00 dansmith or you meant "isn't detected to abort"
20:15:56 sean-k-mooney i think we shoudl abort starting up if the hypervior_hostname for a compute node changes yes
20:16:03 sean-k-mooney at least for libvirt
20:16:30 sean-k-mooney but again thats an extra check we can add to the end fo the series
20:17:15 sean-k-mooney so i proably should have a vm runnign while i do this as i think the instance.node is going to be out of sync
20:17:22 sean-k-mooney so ill boot one before i fix it
20:17:58 sean-k-mooney and my expectation is that when i restor the old hostname the isntance.node will not be updated but placement and the compute nodes table will be
20:18:10 sean-k-mooney thats proably the last test ill do tonight
20:18:18 sean-k-mooney but ill do more on monday.
20:18:25 dansmith yeah, like I said, out of scope of what I'm trying to do, but probably should be in scope
20:18:27 dansmith ack
20:19:23 sean-k-mooney actully i need to test with provider.yaml too at some point
20:19:41 sean-k-mooney we can refernce the RP by name i belive which uses the hypervior_hostname
20:19:48 sean-k-mooney we can also refrrence it by UUID
20:20:05 sean-k-mooney so i suspect the uuid way will be fine but the other way might break
20:20:20 sean-k-mooney again not nessiarly in scope but i want to know what happens
20:41:27 sean-k-mooney ok so if there are allocation in placment we cant delete and rename the resouce provider becaues we get a 409
20:42:08 sean-k-mooney so that fails and we never update the compute node record either
20:42:32 sean-k-mooney the perodic is broken
20:43:24 sean-k-mooney but the agent does not exit and we just get tracebacks in the logs.
20:44:12 sean-k-mooney ill leave it there for today and come back to this on monday
20:46:34 dansmith wait,
20:46:43 dansmith I thought you said it *did* update the placement provider hostname?
20:47:56 sean-k-mooney it appeard too but based on that log it deleted the RP and recreated it the the old uuid and new name

Earlier   Later