| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-03-16 | |||
| 21:19:49 | melwitt | artom: yeah. I already got my ass kicked looking at http://status.openstack.org/elastic-recheck/#1844929 spent days digging in and no dice so far | |
| 21:20:11 | artom | mriedem, wouldn't that depend on the Neutron backend though? | |
| 21:20:17 | mriedem | artom: i think i just said why :) and it's in the bug | |
| 21:20:34 | artom | mriedem, I'm pretty sure at least some of them would wait until libvirt plugs the VIF before sending out the event | |
| 21:20:39 | melwitt | artom: but anecdotally I see http://status.openstack.org/elastic-recheck/#1813789 fail really often in my gerrit notifications so argh ... just want to fix some of these | |
| 21:21:01 | mriedem | we use ovs in the gate and for that backend neutron sends the event when the port binding host changes | |
| 21:21:09 | sean-k-mooney | artom: ovs waits to send the event yes | |
| 21:21:11 | mriedem | which is why we *don't* get the event for things like hard reboot | |
| 21:21:32 | artom | mriedem, sean-k-mooney, get your stuff in line and stop contradicting yourselves ;) | |
| 21:22:16 | artom | melwitt, I get you - it's annoying and you want to fix it | |
| 21:22:22 | artom | It's just such a mess | |
| 21:22:26 | sean-k-mooney | mriedem: do you know if we create a second port binding and activate it our just update the host | |
| 21:22:35 | sean-k-mooney | mriedem: that would cahnge when teh event is sent | |
| 21:22:47 | mriedem | evacuate doesn't use multiple port bindings like live migration | |
| 21:23:07 | mriedem | when i left the only things that used multiple port bindings were live migration and cross-cell resize | |
| 21:23:12 | sean-k-mooney | ok we the condition to sent the event is the port must be in the active state and be bound to a host | |
| 21:23:40 | sean-k-mooney | so since its already in an active state when we bind it in the evacuate it might send the event imideatly | |
| 21:23:41 | artom | sean-k-mooney, so what mriedem was saying then | |
| 21:23:58 | sean-k-mooney | yep i havent checkt it but i would guess that he is correct | |
| 21:24:18 | artom | melwitt, well, if you want to take a whack at it, you could probably use the bind-time stuff I added to the model to change when evacuate starts listening for the event | |
| 21:24:23 | melwitt | ok so the main idea is find a way to register the callback earlier on | |
| 21:24:25 | artom | But there be dragons | |
| 21:24:47 | sean-k-mooney | artom: if we careted a second port binding it would not send it until we activate it | |
| 21:25:49 | melwitt | artom: thanks | |
| 21:26:17 | artom | melwitt, well, yes and no - looking at my own code, depending on whether the port has what I called "bind_time_events", you wait in the compute manager when you send the Neutron request | |
| 21:26:45 | artom | And if they're "plug_time_events", you wait in the virt driver when you plug the VIFs | |
| 21:26:53 | melwitt | ahhh ok | |
| 21:26:54 | sean-k-mooney | mriedem: yes those are still the only things that use the multiple prot bindings | |
| 21:27:04 | artom | melwitt, also, ask sean-k-mooney ;) | |
| 21:27:08 | artom | (Bus, meet Sean :D ) | |
| 21:27:13 | melwitt | lol | |
| 21:27:32 | sean-k-mooney | artom: then you just delete the code and start again | |
| 21:27:46 | artom | Can we do that for all of Nova? ;) | |
| 21:28:16 | melwitt | I think I vaguely get it. I can read though your patch, just knowing a generic idea of what's going on helps a lot. saves a lot of time | |
| 21:29:04 | mriedem | "if we careted a second port binding it would not send it until we activate it" is a bigger non-backportable change most likely because of the behavior changes between compute and conductor | |
| 21:29:06 | artom | melwitt, ping me if you have questions / need review / whatever | |
| 21:29:34 | melwitt | thanks ++ | |
| 21:29:38 | artom | (/me needs pressure to "re-join" upstream) | |
| 21:29:45 | artom | I've been neglecting y'all | |
| 21:30:08 | mriedem | don't forget to loop dansmith into this when you want to talk about it, i'm sure he'd love to | |
| 21:30:38 | melwitt | artom: heh. I might end up running away from this screaming after I try to work on it, so if that doesn't happen maybe I'll ping you | |
| 21:30:38 | artom | You're such a good friend | |
| 21:31:01 | artom | melwitt, screaming would be a good sign, actually | |
| 21:31:02 | sean-k-mooney | ya we cant backport adopting multiple port bindings for evacuate | |
| 21:31:11 | artom | Means you're sane (inasmuch as that's still possible) | |
| 21:31:14 | sean-k-mooney | im not sure if we want to do that or not in general | |
| 21:31:18 | melwitt | lol | |
| 21:31:36 | sean-k-mooney | it might be useful but its alot of work to untangel things and make sure it works | |
| 21:32:14 | artom | Yep. We'd also need to run it with a couple of other Neutron backends | |
| 21:32:28 | artom | IIRC we created a DNM job to run against... OVS? OVN? | |
| 21:32:38 | artom | OVB? OG? | |
| 21:32:39 | sean-k-mooney | ovs ovn and lb | |
| 21:32:47 | artom | RunDMC? | |
| 21:33:20 | melwitt | 😂 | |
| 21:33:20 | melwitt | 😂 | |
| 21:33:30 | melwitt | had to use an emoji for that one | |
| 21:34:00 | sean-k-mooney | sure you "had too" :P | |
| 21:34:24 | melwitt | yeah, it wouldn't let me type anything else until I posted the emoji | |
| 21:35:21 | openstack | Launchpad bug 1813789 in OpenStack Compute (nova) "Evacuate test intermittently fails with network-vif-plugged timeout exception" [Medium,In progress] - Assigned to Artom Lifshitz (notartom) | |
| 21:35:21 | sean-k-mooney | did https://bugs.launchpad.net/nova/+bug/1813789 come up recently downstream or in relation to the nova-livemigation job | |
| 21:35:33 | sean-k-mooney | i rembere talking about it a few days ago | |
| 21:35:39 | sean-k-mooney | i just dont recall the context | |
| 21:35:53 | melwitt | lyarwood mentioned it in the nova meeting | |
| 21:36:15 | sean-k-mooney | ah ya it was in context of the zuul v3 migration | |
| 21:36:18 | melwitt | just saying he hit that bug and http://status.openstack.org/elastic-recheck/#1844929 a bunch of times while trying to get some work done | |
| 21:36:35 | sean-k-mooney | yep | |
| 21:37:08 | melwitt | so I started looking at http://status.openstack.org/elastic-recheck/#1844929 and got nowhere. and now I know http://status.openstack.org/elastic-recheck/#1813789 is also horrid | |
| 21:38:44 | sean-k-mooney | so given we dont use the multiple port bindings flow i think we can assume that as long as the port is still active it will recive a bind time even rather then plug time | |
| 21:38:50 | mriedem | those probably aren't even in the same ballpark of terrible | |
| 21:39:15 | sean-k-mooney | but i would need to think that true more carefully to make sure that is correct | |
| 21:39:46 | mriedem | in the bug i linked in logs where the things were happening so it's not really a question of where the race is | |
| 21:39:57 | mriedem | though those log links are going to be dead by now | |
| 21:40:11 | mriedem | unshelve has the same issue | |
| 21:40:39 | mriedem | 1. bind to new host triggers async network-vif-plugged event, 2. driver.spawn plugs vifs which sets up the callback handler | |
| 21:40:47 | mriedem | if you get the event before 2 you're stuck | |
| 21:40:55 | melwitt | right, ok | |
| 21:41:42 | sean-k-mooney | mriedem: in the unsevle case you could argue that when we go to shevle offloaded teh status of the port shoudl be down | |
| 21:42:08 | sean-k-mooney | which woudl prevent the event being sent in the ovs case atleast until its plugged in the driver | |
| 21:42:13 | mriedem | yup, i opened an old bug for that as well | |
| 21:42:32 | sean-k-mooney | but we would need to sitll use the bind_time vs plugtime thing to check | |
| 21:43:13 | sean-k-mooney | is the state of the port something we can contol from nova? | |
| 21:43:26 | sean-k-mooney | or rather are we allowed to set it | |
| 21:43:38 | sean-k-mooney | if so we could set it to down when we do an evaucate | |
| 21:44:34 | melwitt | well, I'll try looking at it, see how it goes | |
| 21:47:28 | mriedem | this is the thing i was thinking about for a shelve related bug https://github.com/openstack/nova/blob/stable/stein/nova/network/neutronv2/api.py#L3332 | |
| 21:47:55 | mriedem | looks like i updated that as part of the cross-cell series https://github.com/openstack/nova/blob/master/nova/network/neutron.py#L3339 | |
| 21:48:43 | mriedem | https://review.opendev.org/#/c/697162/ | |
| 21:48:46 | mriedem | it's all coming back to me | |
| 21:48:53 | sean-k-mooney | so in that function we would jsut set the binding host to None | |
| 21:49:00 | sean-k-mooney | which would unbind it | |
| 21:49:18 | sean-k-mooney | and the status should go to down as a result | |
| 21:50:19 | mriedem | yeah like how _unbind_ports works | |
| 21:50:38 | mriedem | except you can't clear the device_owner on the port when shelve offloading | |
| 21:50:45 | mriedem | the nova instance needs to continue to "own" the port | |
| 21:51:09 | sean-k-mooney | we need to keep device_id which is the nova instance uuid too but ya | |
| 21:51:43 | mriedem | let us *shelve* this discussion for 6 months from now when it comes up again :) | |
| 21:51:56 | sean-k-mooney | :) | |
| 21:51:57 | mriedem | o/ | |
| 21:57:20 | sean-k-mooney | melwitt: so calling self.network_api.cleanup_instance_network_on_host on the source node durign an evacuate might allow us to use the bind vs plug time evnet code to determin when to wait in the evacuate case too | |
| 21:57:50 | sean-k-mooney | melwitt: it is currently only called for cross cell resize | |