| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-02-02 | |||
| 15:18:30 | bauwser | we're grown-ups, we can do the unplug/plug ourselves | |
| 15:18:34 | bauwser | if Neutron catches it, cool | |
| 15:18:44 | bauwser | but that shouldn't be a blocking situation | |
| 15:18:45 | Spazmotic | what is neutron need to update on the port? Just the state field? | |
| 15:19:02 | Spazmotic | It's too late to english | |
| 15:20:20 | melwitt | bauwser: originally it was just because we wait during instance boot so I thought that's the normal flow is plug and wait for vif-plugged. but then this bug cropped up and then I found the other comment from mriedem in the file and realized there are situations where we should not wait | |
| 15:20:42 | bauwser | melwitt: for the boot case, I agree, it's important | |
| 15:21:02 | mriedem | this does likely mean that after a reboot, we say the instance is active but it's not really, since you can't ssh into it, | |
| 15:21:08 | bauwser | melwitt: for a reboot, the unplug is transient because we want to just cleanup the instance | |
| 15:21:10 | mriedem | which was the problem we had with initial create | |
| 15:21:38 | bauwser | arf, I need to do some parenting for 20-ish mins | |
| 15:21:55 | melwitt | that's still true, fwiw. there's going to be some time until ssh is possible. but I guess you mean not even getting connection refused | |
| 15:22:23 | melwitt | bauwser: I know, just saying why the wait was put there initially | |
| 15:22:42 | melwitt | I don't think it should be there now, based on how it behaves and based on the other comment in the file | |
| 15:22:58 | bauwser | melwitt: yup, I just said in the review that with your patch, the only case now where we wait is boot | |
| 15:23:07 | melwitt | yeah | |
| 15:23:24 | mriedem | what did we do before the "blow everything away on reboot"? did we wait, or we just didn't touch the vifs at all? | |
| 15:23:53 | melwitt | didn't touch the vifs | |
| 15:24:22 | melwitt | well, the code has always called plug_vifs for some reason, but before, we didn't *unplug* them | |
| 15:24:40 | melwitt | so we'd call plug_vifs on vifs that were already plugged | |
| 15:25:01 | melwitt | (since _create_domain_and_network calls plug_vifs unconditionally) | |
| 15:25:55 | openstackgerrit | Matt Riedemann proposed openstack/nova master: libvirt: fix native luks encryption failure to find volume_id https://review.openstack.org/539739 | |
| 15:26:00 | mriedem | yaaaaarwood: melwitt: ^ now handling pre_live_migration also and unfuckified the test | |
| 15:28:21 | melwitt | ack | |
| 15:28:26 | Spazmotic | Alright well. goodnight then folks. | |
| 15:28:59 | openstackgerrit | Ed Leafe proposed openstack/nova master: Escalate UUID validation warning to error in test https://review.openstack.org/540386 | |
| 15:29:17 | figleaf | giblet: ^^ I'm not seeing the exception being raised for bad UUIDs | |
| 15:31:04 | mriedem | melwitt: reading back over the related bug for that comment i left in the resize code, https://bugs.launchpad.net/nova/+bug/1323658 | |
| 15:31:05 | openstack | Launchpad bug 1323658 in neutron "Nova resize/restart results in guest ending up in inconsistent state with Neutron" [Medium,Fix released] - Assigned to Dariusz Smigiel (smigiel-dariusz) | |
| 15:31:12 | mriedem | phil day linked this related change https://review.openstack.org/#/c/68942/ | |
| 15:31:19 | mriedem | to wait for the guest to stop | |
| 15:31:44 | mriedem | if we're blowing away the guest on a reboot, i wonder if we should be doing that also, although if you're doing a soft reboot i think we give the guest os some time and then hard reboot just blows it away | |
| 15:32:04 | melwitt | yeah, our case is a hard reboot | |
| 15:32:35 | melwitt | meaning you're right, the soft reboot does the graceful shutdown thing and hard reboot doesn't | |
| 15:36:29 | melwitt | I'm trying asking in #openstack-neutron | |
| 15:36:38 | sean-k-mooney | mriedem: in the case of ovs the neutron ovs agent spawns a monitor process that recives noticiation every time a new interfaces is added to ovs | |
| 15:37:47 | openstackgerrit | Chris Dent proposed openstack/nova master: Don't rely on parse.urlencode in url comparisons https://review.openstack.org/540420 | |
| 15:38:00 | mriedem | melwitt: this is the closest i get in the history on that other bug for an explanation on why we don't wait for vif plugged on resize https://bugs.launchpad.net/nova/+bug/1323658/comments/57 | |
| 15:38:01 | openstack | Launchpad bug 1323658 in neutron "Nova resize/restart results in guest ending up in inconsistent state with Neutron" [Medium,Fix released] - Assigned to Dariusz Smigiel (smigiel-dariusz) | |
| 15:39:44 | belmoreira | melwitt mriedem we upgraded to Ocata yesterday and I'm observing that "memory_mb_used" in "compute_nodes" tables is not correct. However, the resourcetracker log in the hypervisor is correct. | |
| 15:39:52 | belmoreira | Is there a known bug that you are aware about this issue? | |
| 15:40:16 | melwitt | mriedem: nice sleuthing | |
| 15:41:14 | mriedem | melwitt: so in the resize case, we don't wait for vif-plugged events on the destination host because the vifs are already plugged on the source host, and the source host vifs aren't unplugged until the user confirms the resize and we cleanup the guest from the source | |
| 15:41:16 | melwitt | ihar said that if os-vif doesn't rebind the port, neutron would probably not be aware of the unplug/plug. I'm not sure whether it rebinds it | |
| 15:44:13 | sean-k-mooney | mriedem: so when os-vif adds a port we set the neurtorn port uuid in the external_ids field and wehn it get the notification fo the port added the neutron agent looks up the port and does the binding | |
| 15:44:32 | sean-k-mooney | melwitt: os-vif should not be binding the port port binding and port plugging are 2 different things | |
| 15:44:50 | sean-k-mooney | melwitt: you cannont call os-vif wiht an unbound port | |
| 15:45:07 | melwitt | thanks sean-k-mooney. I really need to read a book about this stuff | |
| 15:45:24 | sean-k-mooney | melwitt: if you do you proably will wish you had not :P | |
| 15:45:30 | melwitt | heh, fair | |
| 15:45:45 | mriedem | sean-k-mooney: it sounds like this behavior is different based on if you're using linuxbridge or not | |
| 15:46:45 | sean-k-mooney | mriedem: kind of. the general patteren is the same but some of the details are different | |
| 15:46:55 | Spazmotic | melwitt, you happen to have time to talk? | |
| 15:47:12 | melwitt | Spazmotic: that sounds mysterious | |
| 15:47:16 | sean-k-mooney | mriedem: for ovs neuton can subscrible to notification form the ovsdb for port add/remove events | |
| 15:47:26 | Spazmotic | Lol nah nah I assure you, just a quick conversation about nova related activities | |
| 15:47:39 | sean-k-mooney | melwitt: for linuxbridge i belive they have to poll | |
| 15:47:47 | sean-k-mooney | * mriedem ^ | |
| 15:47:47 | Spazmotic | If you're too busy it's no rush :) | |
| 15:47:54 | melwitt | Spazmotic: go ahead | |
| 15:48:25 | mriedem | sean-k-mooney: really just trying to figure out if/why vif plug events from neutron might be different in the case of a hard reboot if you're using LB or OVS | |
| 15:48:40 | mriedem | since the wait logic on the nova side would suck if we had to know it changes per network backend | |
| 15:49:18 | openstackgerrit | Eric Fried proposed openstack/nova master: Use update_provider_tree from resource tracker https://review.openstack.org/520246 | |
| 15:49:18 | openstackgerrit | Eric Fried proposed openstack/nova master: Fix nits in update_provider_tree series https://review.openstack.org/531260 | |
| 15:49:19 | openstackgerrit | Eric Fried proposed openstack/nova master: Move refresh time from report client to prov tree https://review.openstack.org/535517 | |
| 15:49:19 | openstackgerrit | Eric Fried proposed openstack/nova master: Make generation optional in ProviderTree https://review.openstack.org/539324 | |
| 15:50:08 | sean-k-mooney | mriedem: if you do a hard reboot we will actully terminate the qemu process and remove the tap and re add it | |
| 15:50:23 | sean-k-mooney | if you are polling you may miss that | |
| 15:50:50 | sean-k-mooney | with ovs you will get the notifcation from the ovsdb for the add and remove | |
| 15:52:19 | sean-k-mooney | mriedem: is this all in relation to https://bugs.launchpad.net/nova/+bug/1323658 | |
| 15:52:20 | openstack | Launchpad bug 1323658 in neutron "Nova resize/restart results in guest ending up in inconsistent state with Neutron" [Medium,Fix released] - Assigned to Dariusz Smigiel (smigiel-dariusz) | |
| 15:52:36 | mriedem | sean-k-mooney: no, https://review.openstack.org/#/c/540168 | |
| 15:57:16 | bauwser | damn, interesting convo about vif pluging and neutron agents | |
| 15:57:24 | giblet | figleaf: Locally it raises for me if I pull the oslo ovo patch into my env (Depends-On: https://review.openstack.org/540344) | |
| 15:57:38 | giblet | figleaf: I mean it raised before the rebase | |
| 15:57:55 | giblet | figleaf: let me try with the new patch | |
| 15:58:22 | bauwser | mriedem: yeah, I guess the crux of the problem is how linuxbridge notifies the events | |
| 15:58:34 | bauwser | since that bug report is all about it | |
| 15:59:07 | bauwser | I mean for Quantum | |
| 15:59:08 | sean-k-mooney | welll so i think we are conflating too things here | |
| 15:59:16 | sean-k-mooney | vif pluging and vif binding | |
| 15:59:27 | bauwser | vif binding is the allocation, right? | |
| 15:59:28 | sean-k-mooney | vif plugging is the act of adding the tap to the backend | |
| 15:59:40 | bauwser | vif plugging is the fact to attach the vif to the guest ? | |
| 16:00:03 | sean-k-mooney | vif binding id that ack of the neutron ml2 direver populating the port details with a vif_type and vif_binding_details | |
| 16:00:25 | mriedem | this might be closer to the issue https://review.openstack.org/#/c/540168/1/nova/virt/libvirt/driver.py@5372 | |
| 16:00:38 | mriedem | for one, we're operating on a stale cache to determine which vifs to wait for | |
| 16:00:46 | mriedem | b/c we never wait for the vifs to be unplugged and refresh the nw info cache | |
| 16:01:02 | mriedem | based on that, it's likely better to yes just ignore the vif plugged event, | |
| 16:01:19 | mriedem | in the server create case, we have the async network info wrapper that is waiting for the nw_info cache updates during port binding, | |
| 16:01:24 | mriedem | and then the virt driver spawn is waiting on those | |
| 16:01:40 | mriedem | so the nw info cache is up to date in the server create case, but we have nothing like that in the case of hard reboot | |
| 16:02:26 | melwitt | yeah, IIUC we get the network_info fresh at the beginning of a reboot and then that's it | |
| 16:02:33 | sean-k-mooney | mriedem: well one thing im not clear on. when we do a hard reboot do we redo the neutron port binding or jsut reuse the old one | |
| 16:02:43 | mriedem | re-use the existing one | |
| 16:03:38 | sean-k-mooney | well the neutron event is sent when neutron finish wirieing up a prot for the first time after portbinding. i dont know if neutron will send it again if we dont rebind | |
| 16:04:06 | mriedem | it must, otherwise this would be 100% fail in the gate | |
| 16:04:09 | mriedem | what we have is a race right now | |
| 16:04:25 | sean-k-mooney | mriedem: oh thats good to know | |
| 16:05:37 | mriedem | melwitt: i think we should probably not copy the comment from finish_migration | |