Earlier  
Posted Nick Remark
#openstack-nova - 2018-02-16
21:20:57 fried_rice Exacitically
21:21:08 fried_rice L518-9 in fact
21:21:20 jroll right
21:23:02 jroll we demonstrate that code is running: http://logs.openstack.org/50/544750/10/check/ironic-grenade-dsvm-multinode-multitenant/d7a1ee7/logs/subnode-2/screen-n-cpu.txt.gz#_Feb_16_17_21_04_613786
21:24:05 fried_rice I'm sorta guessing you can't change the UUID of a ComputeNode in the db. mriedem?
21:24:14 jroll both compute nodes appear to be attempting to create an RP for that node (9de4d7b4-51c9-4088-99b4-cd648332504e), at different times of course
21:24:31 fried_rice Right; the first one succeeds but the second one barfs
21:24:33 fried_rice right?
21:24:34 jroll um
21:24:43 jroll well, both fail
21:24:48 mriedem fried_rice: changing the uuid would be kinda bad
21:24:52 jroll but my um: how do we feel about adding a cn.save() here: https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py?utf8=%E2%9C%93#L525
21:25:11 fried_rice jroll: I was going to suggest that.
21:25:20 fried_rice but also not knowing if it's The Right Thing.
21:25:28 fried_rice I was also going to suggest setting the UUID before save()ing.
21:25:35 fried_rice But mriedem won't come to my birthday party if I do that.
21:25:51 jroll nah, the UUID should be fine, we're updating an existing CN
21:26:21 fried_rice But I thought that was the whole problem.
21:26:46 fried_rice btw, assuming _resource_change returns True, that .save() should be getting done at L859.
21:27:28 jroll yeah, I'm not sure it will, the schedulable resources should be the same as before, only the host attribute is changing
21:30:14 jroll ok, I get it now, we did get a new UUID, wtf
21:30:29 fried_rice w, indeed, tf
21:30:47 jroll sorry, I got confuzzled
21:31:40 jroll you know...
21:32:04 jroll does a compute_node record get deleted at some point if the compute service disappears?
21:32:13 jroll and if so, do we clean up the resource providers?
21:33:09 fried_rice I could answer that second question if I knew the answer to the first. But I don't. mriedem? (Or mriedem if you'd care to delegate, who's around who knows this stuff?)
21:33:36 jroll so um
21:33:43 jroll this is pike->queens, afaik
21:33:53 jroll can we please land https://review.openstack.org/#/c/527423/ .
21:33:56 openstackgerrit Eric Fried proposed openstack/nova master: WIP: Make sure rebalance saves the compute node https://review.openstack.org/545464
21:34:07 fried_rice jroll: Here's a quickie to make sure that save() is happening ^
21:34:40 fried_rice jroll: But where the error happens, we're clearly running that code.
21:34:45 jroll er, the one I'm looking at is queens->master
21:34:47 jroll yeah
21:35:01 fried_rice So whereas I agree we should land that backport, that's not gonna be your fix here.
21:35:15 mriedem jroll: it doesn't
21:35:20 jroll sorry, we've been having issues with this job since before queens was cut, so I've been getting confused
21:35:23 mriedem compute_nodes hang out until manually removed
21:35:29 jroll damn
21:35:56 mriedem there are some bugs that tssurya opened for removing compute nodes and providers
21:36:14 jroll I was hoping compute nodes did get deleted but not RPs
21:36:18 jroll that would explain things
21:36:26 mriedem https://bugs.launchpad.net/nova/+bug/1749734
21:36:26 openstack Launchpad bug 1749734 in OpenStack Compute (nova) "Purge the compute_node records, resource provider records and host_mappings when doing force delete of the host" [Medium,Confirmed] - Assigned to Surya Seetharaman (tssurya)
21:36:30 jroll thanks
21:37:26 fried_rice jroll: I can get a little more aggressive with the hacking.
21:38:27 jroll fried_rice: feel free, I'm just trying to wrap my head around some of this
21:38:48 jroll this seems wrong. http://logs.openstack.org/50/544750/10/check/ironic-grenade-dsvm-multinode-multitenant/d7a1ee7/logs/screen-n-cpu.txt.gz#_Feb_16_17_06_33_592313
21:39:09 jroll oh, there we are http://logs.openstack.org/50/544750/10/check/ironic-grenade-dsvm-multinode-multitenant/d7a1ee7/logs/screen-n-cpu.txt.gz#_Feb_16_17_04_22_814020
21:39:35 jroll but we can't talk to placement, so the RP stays
21:39:57 jroll because Feb 16 17:04:22.591129 ubuntu-xenial-rax-ord-0002580076 nova-compute[28778]: DEBUG nova.virt.ironic.driver [None req-969cdb75-026b-4cac-ba09-3f3be962a09d service nova] Returning 0 available node(s) {{(pid=28778) get_available_nodes /opt/stack/old/nova/nova/virt/ironic/driver.py:757}}
21:40:01 jroll because ironic is down
21:40:05 jroll got dang.
21:40:30 jroll I remember this code landing to fix something else
21:40:47 fried_rice jroll: Is there any chance that this ironic node being rebalanced has allocations?
21:41:05 jroll fried_rice: yes, we create an instance before the upgrade AFAIK
21:41:43 fried_rice So we have to do more than just delete the old RP. We have to move his allocations to the new one. This ain't gonna work.
21:41:50 jroll this crap is burning us: https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L607-L616
21:41:57 jroll right, we shouldn't be deleting the compute node or the RP
21:42:16 jroll so here's what's going on, in short:
21:42:26 fried_rice The answer is quite simply to make sure the compute node doesn't change its freakin uuid.
21:42:42 jroll n-cpu has a bunch of ironic nodes it's managing
21:42:48 jroll ironic goes down for upgrade
21:42:54 jroll n-cpu does a RT update
21:42:59 jroll n-cpu can't reach ironic
21:43:18 jroll n-cpu thinks all the ironic nodes are gone, for good, as if ironic returned I have no nodes
21:43:25 jroll n-cpu deletes the compute_node records
21:43:28 jroll ironic comes back
21:43:40 jroll n-cpu does an RT update, sees nodes, creates compute_node records
21:44:17 jroll meanwhile, n-cpu couldn't delete the resource providers from placement, and so it tries to create new ones and *boom*
21:45:02 fried_rice n-cpu deletes the compute_node records?
21:45:02 jroll btw, s/ironic goes down for upgrade/keystone goes down for upgrade/, which is why neither ironic nor placement can be reached
21:45:17 fried_rice I thought we decided that wasn't happening.
21:45:21 jroll yes
21:45:25 jroll http://logs.openstack.org/50/544750/10/check/ironic-grenade-dsvm-multinode-multitenant/d7a1ee7/logs/screen-n-cpu.txt.gz#_Feb_16_17_04_22_814020
21:45:40 jroll because: http://logs.openstack.org/50/544750/10/check/ironic-grenade-dsvm-multinode-multitenant/d7a1ee7/logs/screen-n-cpu.txt.gz#_Feb_16_17_04_22_591129
21:45:48 jroll because: https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L607-L616
21:46:42 jroll (maybe it is apache that is down, don't know, don't care, things are down and n-cpu can't deal)
21:47:44 jroll for the curious, the commit message and bug here explain why we're returning an empty list of nodes there: https://review.openstack.org/#/c/487925/
21:48:04 jroll fried_rice: that all make sense?
21:48:11 jroll TheJulia: ^ fyi, I think I nailed it down.
21:50:33 jroll thanks, my gas can is nearly empty
21:53:33 jroll turns out lying to the resource tracker is wrong, who'da thought
21:57:05 jroll so the only way n-cpu will blow up by ironic not being reachable is because of this: https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L524
21:57:10 jroll which will happen on the first RT run
21:57:54 jroll or not even? wtf
21:58:29 jroll oh, there used to be a _refresh_cache() there
21:58:54 jroll but even before 487925 we would return an empty list
21:59:57 jroll ah jeez https://github.com/openstack/nova/commit/cce06a1e9855d9eed3f7c653200853f23466d791
22:01:19 jroll hm, 5pm friday
22:01:35 fried_rice jroll: Need anything from me?
22:01:52 jroll fried_rice: whiskey may be needed
22:01:54 jroll :)
22:01:56 fried_rice Can https://review.openstack.org/545464 be abandoned?
22:02:06 jroll yes, believe so
22:02:27 fried_rice jroll: The fact that you're a time zone ahead of me indicates I have no way of getting you a bottle in time to save you.
22:02:27 TheJulia jroll: I will buy you whiskey in Dublin
22:02:34 fried_rice Yeah, that ^
22:02:36 jroll heh

Earlier   Later