Earlier  
Posted Nick Remark
#openstack-nova - 2018-02-02
16:11:46 sean-k-mooney mriedem: if that arge is set to True we will not wait for neutron
16:12:00 bauwser sean-k-mooney: well, all other server actions do the same, y'kno<
16:12:05 bauwser know*
16:12:14 sean-k-mooney os-vif will add the port when we call plug
16:12:32 bauwser sean-k-mooney: we're only able to correctly wait for Neutron in the boot case only because we have a fresh network info
16:12:37 sean-k-mooney but we wait for neutron to make sure the firewal rules are active on the port before unpasing the vm
16:12:45 bauwser fixing that in other server action calls would require far more
16:13:04 bauwser we == os-vif ?
16:13:15 sean-k-mooney we == nova
16:13:29 melwitt mriedem: he's the run I was looking at yesterday. there are no network-vif-unplugged events for the instance for the failed test (the hard reboot) http://logs.openstack.org/42/525842/11/check/neutron-tempest-linuxbridge/2502b64/logs/screen-q-svc.txt.gz#_Jan_29_16_01_28_170964
16:13:30 sean-k-mooney os-vif does not talk to neutron ever
16:13:53 melwitt I linked to a network-changed but if you look further in the log there are no unplug events
16:14:14 melwitt so I don't expect things would work if we tried to wait for an unplug event because neutron never sends one
16:16:14 sean-k-mooney melwitt: correct but wen we delete the tap any iprules create by the neutron security group driver are removed. if neutron never sees the device is unplugged and plugged again i dont think it will ever recreate them
16:18:18 melwitt sean-k-mooney: os-vif unplug deletes the tap, is what you are saying?
16:18:22 mriedem giblet: thanks for reviewing the nova-multiattach job change, replies inline https://review.openstack.org/#/c/532689/
16:18:51 sean-k-mooney melwitt: actully qemu doing a hard reboot does but os-vif will also remove it form the linux bridge
16:19:06 sean-k-mooney or ovs bridge depending on your life choices
16:19:32 melwitt sean-k-mooney: okay. and then the os-vif plug call will create *not* create the iprules again?
16:19:57 sean-k-mooney melwitt: the ip routes are created by the neutron security group driver not os-vif
16:20:19 melwitt hm
16:20:25 giblet mriedem: cool. looking in a minute
16:20:35 sean-k-mooney so if nova does not notify neutron i dont think it will reinstall them
16:20:57 sean-k-mooney melwitt: we may be able to simply rebind the port to the same host as part of the reboot
16:21:29 melwitt sean-k-mooney: if that's the case, I think that would be a latent bug because the behavior before this "re-initialize everything" approach still destroyed and undefined the libvirt domain. it just didn't do any os-vif unplug/plug
16:21:36 mriedem should we just not destroy_vifs on hard reboot and avoid this mess?
16:22:27 bauwser mriedem: well, looks like we somehow decided in the past to destroy everything
16:22:33 sean-k-mooney mriedem: the tap will still be delete by qemu/kvm/libvirt so its not an os-vif issue
16:22:49 mriedem bauwser: decided in the last 6 months
16:22:58 mriedem bauwser: this is a regression introduced in queens,
16:22:59 bauwser interesting
16:23:06 mriedem which is why i'm -W on the ocata backport
16:23:11 melwitt no, the handling of the libvirt domain did not change
16:23:14 openstackgerrit Balazs Gibizer proposed openstack/nova master: Escalate UUID validation warning to error in test https://review.openstack.org/540386
16:23:31 giblet figleaf: now your new unit test also passes in my local env ^^
16:23:50 giblet figleaf: btw thanks for that unit test. I did not now that we have tests for the fixtures
16:25:35 sean-k-mooney so i think we dont see this on ovs because of the ovsdb monitor which will see the port remved and readded by libvirt/os-vif. the gap seams to be that we dont have a reliable way to notify neutron that this happened for linuxbridge
16:26:35 melwitt if we're losing iptables rules upon reboot, we've always been losing them, if it has to do with destroying the libvirt domain because we have always been destroying the domain on a reboot and we also destroy it for a stop, etc
16:27:34 bauwser yeah, stop/start calls hard_reboot
16:27:35 bauwser fun
16:27:36 sean-k-mooney mlavalle: but only with linuxbridge when the polling intergave misses the removal and addtion
16:28:43 sean-k-mooney /intergave/interval/ that was bad even for me
16:29:04 melwitt my point is that is not new, that was not changed in the last 6 months. the change was doing os-vif unplug/plug and volume detach/attach
16:29:38 mriedem melwitt: i think you pointed this out before but this is likely a bogus debug message https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L346
16:29:49 mriedem _events being empty could just mean we already processed the event
16:29:53 sean-k-mooney melwitt: oh ok
16:29:58 mriedem _events == None means we are shutting down compute
16:30:20 bauwser melwitt: ping me when you're done with your new rev and I'll review it when I'm back
16:31:36 sean-k-mooney mriedem: you are correct the os-vif is new but the code in the linuxbidge pluging was copied directly from nova so i think the logic should be the same.
16:31:46 sean-k-mooney * melwitt ^
16:32:55 figleaf giblet: We shouldn't have to use the WarningsFixture, as the TestCase class already uses it
16:32:57 melwitt sean-k-mooney: what I was trying to say is that we didn't do anything new about the libvirt domain part in the last 6 months. we have always done the equivalent of a 'virsh destroy' when we do a hard reboot or stop/start
16:33:11 figleaf giblet: Here's what I get when I runt the test: http://paste.openstack.org/show/658860/
16:33:24 melwitt so if that messes up iptables rules, then they've always been getting messed up and it's latent
16:34:00 sean-k-mooney melwitt: oh i know. and we should be doing the equivelent of virsh destroy that is the main difference between soft reboot and hard as far as i am concurened
16:34:23 giblet figleaf: testtools.TestCase does not use the nova's warning fixture just nova.test.TestCase uses it
16:34:39 sean-k-mooney melwitt: yes that is also true
16:35:05 melwitt sean-k-mooney: should or shouldn't? sorry I'm confused. the 'virsh destroy' will indeed wipe out iptables rules?
16:35:32 giblet mriedem: responeded in https://review.openstack.org/#/c/532689
16:35:48 sean-k-mooney a hard reboot should entirely tear down the vm the same way virsh destroy would
16:35:56 melwitt if it does, then I'm surprised we haven't heard about people complaining
16:36:05 giblet mriedem: I'm OK to +W this and continue discussing the run.yaml issue separately as it seems the job works
16:36:32 giblet mriedem: and the rest of the nits can be handled in a follow up (like naming)
16:36:41 sean-k-mooney the soft reboot should jsut send an achi reboot command to the guest but not destry and recreate the entire vm
16:37:12 giblet figleaf: have you tried runnig PS4 in your env or just PS3?
16:37:28 mriedem giblet: works for me
16:37:28 melwitt mriedem: I'm trying to understand how we even get that "Unexpected attempt to pop events during shutdown" in that test run
16:37:50 mriedem melwitt: looking at the n-cpu log, it looks like because we get a network-changed event before that which we process and pop the event, but not sure
16:37:54 melwitt sean-k-mooney: yeah, it does. just trying to understand whether any of that blows away iptables rules
16:38:05 melwitt since that would be bad
16:38:48 sean-k-mooney melwitt: if qemu removes the tap. then any ip tables rules associated with the tap will be removed for the kernel when the tap is removed
16:38:57 mriedem but network-changed shouldn't mess with the _events stuff at all...
16:39:00 giblet mriedem: OK. I've +Wd it
16:39:15 mriedem melwitt: we're probably just getting a network-vif-plugged event that we weren't expecting, so there was nothing in _events
16:39:17 melwitt sean-k-mooney: okay. I don't know what makes qemu decide to remove a tap
16:39:30 mriedem giblet: thanks
16:39:58 ildikov mriedem: giblet: thanks
16:39:59 melwitt mriedem: we did get an unexpected one. we weirdly get network-change AND network-vif-plugged during the hard reboot while we are 'virsh destroy'ing and BEFORE we unplug anything. I don't understand it
16:40:12 mriedem melwitt: yeah was looking at that,
16:40:22 mriedem and after that, the nw info cache says the port is active=false
16:40:28 melwitt right
16:40:30 mriedem so that's why we wait for vif-plugged
16:40:34 mriedem but we aren't going to get that,
16:40:35 melwitt which is how we go into waiting mode
16:40:39 mriedem because we alreayd got a vif-plugged
16:40:46 mriedem yeah...
16:41:00 melwitt I don't get how we get a vif plugged from ... not plugging anything yet
16:41:30 melwitt because you can see the first vif-plugged from the instance boot, then after the reboot starts but before even the unplug, we get vif-plugged again
16:41:34 sean-k-mooney melwitt: well unplug for linux bridge in os-vif is a noop https://github.com/openstack/os-vif/blob/stable/queens/vif_plug_linux_bridge/linux_bridge.py#L114-L117 which is why i dont that that os-vif is the issue here.
16:42:13 melwitt hah, a no-op
16:42:44 mriedem so after the instance is created but before we reboot, we get network-changed here:
16:42:45 mriedem http://logs.openstack.org/42/525842/11/check/neutron-tempest-linuxbridge/2502b64/logs/screen-n-cpu.txt.gz#_Jan_29_16_01_35_168999
16:42:51 mriedem which forces us to refresh the nw info cache
16:43:10 melwitt ah, so that's where the active=True -> active=False change comes from
16:43:21 melwitt I was wondering how that happens
16:43:23 figleaf giblet: no, I was testing on PS3
16:43:30 mriedem then we start rebooting http://logs.openstack.org/42/525842/11/check/neutron-tempest-linuxbridge/2502b64/logs/screen-n-cpu.txt.gz#_Jan_29_16_01_35_600736
16:44:00 figleaf giblet: if you need to explicitly add that to every test class in order to catch invalid UUIDs, then I don't think that's very useful
16:44:07 mriedem "active": true
16:44:07 mriedem at that point, it says the vif is active

Earlier   Later