Earlier  
Posted Nick Remark
#openstack-nova - 2018-02-15
18:49:32 TheJulia http://logs.openstack.org/50/544750/8/check/ironic-grenade-dsvm-multinode-multitenant/5713fb8/logs/screen-n-cpu.txt.gz#_Feb_15_18_00_31_112744
18:49:48 efried lbragstad: Thanks.
18:50:12 lbragstad efried yep
18:50:44 dansmith TheJulia: okay that almost maybe kinda looks like someone changed ids or there's some confusion going on
18:50:58 dansmith TheJulia: multiple nova-computes? multiple ironic nodes?
18:52:14 dansmith actually, efried ^
18:52:15 TheJulia dansmith: 2x n-cpu running pike, 2x ironic-conductor (1 master, 1 queens) 1x ironic api running queens.
18:52:30 jroll and multiple ironic nodes
18:52:35 TheJulia yup
18:52:38 TheJulia 7 of them
18:52:40 dansmith efried: this is during a upgrade, any chance the report client being providertreeish after the upgrade is breaking something?
18:52:48 dansmith because it's ensuring that the provider exists, and then freaking out because it does
18:52:54 jroll reminder it could be a real incompatibility between pike nova and queens ironic, given this test has been off for weeks
18:52:58 jroll dansmith: still pike nova
18:53:08 dansmith jroll: ah, right
18:53:30 TheJulia jroll: it should presently be queens -> master, I'll double check (since the devstack-gate patch was involved)
18:53:47 jroll oh, right
18:53:53 dansmith okay so queens nova code yeah?
18:54:03 jroll which shouldn't be much different than queens/queens which has been working
18:54:06 jroll dansmith: yes, sorry
18:54:12 jroll queens nova on both sides
18:54:15 dansmith ack
18:54:38 TheJulia jroll: we should propably cherry-pick the patch to stable/queens and let it run to see what the case is there...
18:54:57 jaypipes mriedem: heh
18:54:58 jroll TheJulia: wouldn't hurt
18:55:11 TheJulia jroll: going to click button
18:55:15 jroll thank you
18:55:49 dansmith well, queens code still has the provider tree stuff so still worth efried looking at I think
18:56:06 dansmith although weird that it would be different/broken across such a small upgrade boundary
18:56:14 jroll yeah, that's my thought
18:56:15 TheJulia dansmith: in the mean time, we can check the pike -> queens job results and see if it is broken the same way
18:56:18 efried Yup, I'm looking. We shouldn't be able to get here except by really bad timing.
18:56:29 efried TheJulia: And you said it happened more than once?
18:56:39 dansmith more than once in the logs even
18:56:49 dansmith on each sync
18:56:58 jroll could also be a problem with our multinode, whether it's queens+queens or queens+master
18:57:00 TheJulia efried: we have _not_ tried to recheck in case it was some process fluke
18:57:22 efried No, this simply shouldn't happen, still looking...
19:00:12 efried Okay, this is a *name* conflict, not a *uuid* conflict. This means we somehow got two RPs with different UUIDs but the same name.
19:00:26 efried and the name is a UUID, which is nice and confusing.
19:00:35 cfriesen mriedem: with respect to https://review.openstack.org/#/c/544748/ and my proposed change https://review.openstack.org/#/c/525253/. In our case it passes scheduling but then fails for whatever reason on the compute node. As such, the proposed fix is not sufficient because we would not end up calling _bury_in_cell0().
19:00:36 efried I thought johnthetubaguy was banging his head against this last Fall.
19:00:46 jroll efried: the name would be the ironic uuid, right?
19:01:01 efried jroll: I'm not an expert there, but yeah, something like that.
19:01:08 dansmith efried: I wonder if we detected the ironic node go away and come back and we're trying to recreate it but never deleted it?
19:01:14 dansmith jroll: yes
19:01:32 jroll I seem to remember this being john's head banging https://review.openstack.org/#/c/508555/
19:02:11 jroll but there shouldn't be a rebalance happening
19:02:46 efried jroll: Beat me to it, yeah, that's the patch I was thinking of.
19:03:01 efried but there was more to it than that, I thought.
19:03:32 jroll not sure
19:04:15 openstackgerrit Jay Pipes proposed openstack/nova-specs master: mirror nova host aggregates to placement API https://review.openstack.org/545057
19:04:22 jroll the title on the bug certainly looks like this problem: https://bugs.launchpad.net/nova/+bug/1714248
19:04:24 openstack Launchpad bug 1714248 in OpenStack Compute (nova) pike "Compute node HA for ironic doesn't work due to the name duplication of Resource Provider " [High,In progress] - Assigned to Matt Riedemann (mriedem)
19:04:33 efried There was another patch where we mucked with one of the rt update methods to look for things with different names.
19:04:47 dansmith jroll: if ironic was down and the two nova computes notice it is back at different times maybe they could be disagreeing briefly on who owns what?
19:05:11 jroll dansmith: yeah, trying to track down if it's long enough to trigger a rebalance
19:05:58 jroll servicegroup api tells us
19:06:02 melwitt mriedem: FYI, digging into the difference between the cleanup volumes vs ports bugs today
19:06:18 melwitt based on your comments in the patch
19:08:50 openstackgerrit Lee Yarwood proposed openstack/nova stable/queens: DNM: Test LM with encrypted volumes https://review.openstack.org/545093
19:09:38 jroll dansmith: ah yep, the subnode n-cpu considers itself down at this point, I believe http://logs.openstack.org/50/544750/8/check/ironic-grenade-dsvm-multinode-multitenant/5713fb8/logs/subnode-2/screen-n-cpu.txt.gz#_Feb_15_17_59_05_738069
19:10:06 jroll ironic is unreachable for like 5 minutes
19:10:22 jroll or rather a full resource tracker run and then some
19:11:19 openstackgerrit Jackie Truong proposed openstack/python-novaclient master: Microversion 2.61 - Add trusted_image_certificates https://review.openstack.org/500396
19:12:05 mriedem cfriesen: then what you have here https://review.openstack.org/#/c/525253/1/nova/conductor/manager.py doesn't help you
19:12:10 mriedem cfriesen: so i'm confused
19:12:42 mriedem cfriesen: is https://review.openstack.org/#/c/528385/ what you are looking for?
19:12:56 dansmith jroll: yeah, but with johnthetubaguy's reuse-compute-node patch I would think this wouldn't be a problem right?
19:12:58 jroll aaaand we have some problems deleting RPs: http://logs.openstack.org/50/544750/8/check/ironic-grenade-dsvm-multinode-multitenant/5713fb8/logs/screen-n-cpu.txt.gz#_Feb_15_17_58_22_418875
19:13:00 jroll (wtf)
19:13:22 jroll dansmith: I would think so too, just confirming there is likely a rebalance, so they could be disagreeing
19:13:32 dansmith yeah
19:13:50 dansmith jroll: ah, that 503 during deleting is weird
19:14:11 jroll dansmith: indeed
19:14:20 dansmith jroll: do you guys have to restart apache?
19:14:39 TheJulia dansmith: we do
19:14:44 dansmith okay
19:14:45 dansmith also
19:14:47 TheJulia we update the configuration to load a vhost
19:15:02 TheJulia we also shutdown services at 17:55 for the upgrade, nova would have remained running
19:15:09 dansmith if placement is crashing the same way as conductor, maybe placement is dead under apache, hence the 503?
19:15:12 jroll good lord keystone db migrations are spammy
19:15:13 dansmith until you restart?
19:15:20 TheJulia so 17:58 is when everything is down
19:15:31 dansmith thus nova never got to delete the RPs
19:15:55 TheJulia I think the restart is during the ironic upgrade, checking to see when it actually occured
19:16:09 dansmith right, but if placement started crashing during the upgrade of packages,
19:16:21 dansmith which is when nova would have noticed ironic went away and tried to delete RPs or something,
19:16:24 jroll OH
19:16:28 jroll keystone is upgrading there
19:16:35 jroll and so placement can't validate the token
19:16:36 dansmith and then you restart apache..
19:16:38 dansmith ohh
19:16:43 jroll http://logs.openstack.org/50/544750/8/check/ironic-grenade-dsvm-multinode-multitenant/5713fb8/logs/screen-placement-api.txt.gz#_Feb_15_17_58_22_463228
19:17:02 mriedem melwitt: i'm likely also going to start writing a functional test for the case that we delete a build request for a bfv instance during local delete, because we aren't cleaning up volumes there either
19:17:08 dansmith and you get a bug, and you get a bug, and you get a bug...
19:17:21 dansmith jroll: good catch
19:17:28 melwitt mriedem: ack
19:17:50 TheJulia http://logs.openstack.org/50/544750/8/check/ironic-grenade-dsvm-multinode-multitenant/5713fb8/logs/grenade.sh.txt.gz#_2018-02-15_18_00_28_322 is when we restart apache

Earlier   Later