Earlier  
Posted Nick Remark
#openstack-nova - 2018-08-09
20:45:46 sean-k-mooney or rathar the lb l2 agent
20:45:49 dansmith mriedem: so you should be eating them up there, and then his wait is definitely never going to get them right?
20:46:10 dansmith so in pike, I expect it races,
20:46:16 dansmith and in rocky it never ever works at all
20:47:03 mriedem yeah maybe
20:47:10 mriedem i could dig up my 2 node lb ci patch,
20:47:12 dansmith oh sweet baby jesus thank you for this day
20:47:17 mriedem and enable this waiter in nova on master,
20:47:40 mriedem and we'd have to probably turn the vif plugging timeout way down to actually see if we hit a timeout
20:47:53 mriedem otherwise i'd expect in the gate, live migration with a tiny cirros guest not doing anything transfers pretty fast
20:48:03 dansmith yeah
20:48:17 mriedem heh, and i was just going to start mowing and packing
20:48:37 dansmith melwitt: so, honestly, reverting sahid's thing for rocky needs to be high prio I think
20:48:47 dansmith melwitt: live migration with LB is completely broken I expect
20:48:55 mriedem i'll update that ci patch
20:49:02 melwitt ok, so is this a RC1 thing or a RC2 thing?
20:49:06 dansmith melwitt: and it probably needs to be reverted out of the older releases too
20:49:13 dansmith melwitt: your call but RCsomething, IMHO
20:49:36 prometheanfire back
20:49:41 prometheanfire but leaving soonish
20:49:42 dansmith I would think we could do something like what mriedem did in rocky for those releases
20:50:05 melwitt ok, definitely RC2. trying to do it by the end of today will be hard unless it gets approved, stat
20:50:20 melwitt *definitely RC2, at least
20:50:54 dansmith well, maybe not since he needed a signal from the remote machine that it was going to do the wait ...
20:51:02 prometheanfire mriedem: I'll try to get a bug reported if you think that's the next step
20:51:04 dansmith although the event should still trigger
20:51:07 mriedem prometheanfire: yes please
20:51:09 dansmith prometheanfire: yes
20:51:10 mriedem we'll track that for rc2
20:51:27 prometheanfire rc2?
20:51:29 mriedem was going to try and see how long a guest transfer takes in the gate
20:51:37 mriedem prometheanfire: today is release candidate 1 day
20:51:43 prometheanfire I thought this didn't hit rocky
20:51:50 prometheanfire but I'll leave that to you
20:51:53 mriedem prometheanfire: new wrinkle
20:51:57 prometheanfire oh, nice
20:52:00 dansmith worse wrinkle
20:52:02 mriedem (1) probably race fail on pike
20:52:05 prometheanfire happy to help :P
20:52:06 mriedem (2) totes broken on master
20:52:13 prometheanfire even better
20:53:40 dansmith prometheanfire: hook me up with a bug number and I'll propose the revert
20:53:41 dansmith and I can comment on the bug with all the deets
20:53:50 dansmith since mriedem will be busy with the ci patch and packing for da nang
20:53:59 mriedem isn't da nang vietnam?
20:54:02 dansmith yes
20:54:18 openstackgerrit Matt Riedemann proposed openstack/nova master: placement: ignore policy scope check failures if not enforcing scope https://review.openstack.org/590445
20:54:33 prometheanfire I'm hoping to get the user to report, but I will if he went home
20:55:16 dansmith the revert is complete conflict
20:55:19 dansmith wonderful.
20:56:33 dansmith the other benefit of doing this in the manager is that we don't need the silly artificial speed limit
20:56:48 dansmith although we probably need sahid and libvirt people to confirm that there's not something we're missing here
20:57:32 dansmith because I asked him specifically about doing this early on in his patch and he said it wasn't possible, but I believed that qemu/libvirt on the dest machine were responsible for the plugging at the time
20:57:34 dansmith and maybe he did too
21:00:46 mriedem meanwhile, our granite seller is being an ass and i have to get back to our plumber
21:01:37 dansmith #firstworldrichpersonproblems
21:06:47 mriedem ok so looking at a job, we register waiting for events starting here http://logs.openstack.org/98/587498/1/check/nova-live-migration/5ff805a/logs/screen-n-cpu.txt.gz#_Jul_31_17_18_19_562999
21:06:52 mriedem Jul 31 17:18:19.562999 ubuntu-xenial-inap-mtl01-0001077052 nova-compute[2320]: DEBUG nova.compute.manager [None req-942f438c-3cbb-4ce7-8afb-ecd250c98f75 tempest-LiveAutoBlockMigrationV225Test-1515676049 tempest-LiveAutoBlockMigrationV225Test-1515676049] [instance: 7f68c430-f565-433a-8f87-27b9a00d29a0] Preparing to wait for external event network-vif-plugged-6b030652-5fe6-471a-b7ec-0b70e95159a4 {{(pid=2320) prepare_for_instanc
21:06:52 mriedem ent /opt/stack/new/nova/nova/compute/manager.py:328}}
21:07:51 mriedem pre_live_migration takes about 7 seconds
21:07:51 mriedem Jul 31 17:18:26.178239 ubuntu-xenial-inap-mtl01-0001077052 nova-compute[2320]: INFO nova.compute.manager [None req-994d9893-545f-47b9-b93e-1e21cb439db7 tempest-LiveMigrationTest-233418614 tempest-LiveMigrationTest-233418614] [instance: 7f68c430-f565-433a-8f87-27b9a00d29a0] Took 6.61 seconds for pre_live_migration on destination host ubuntu-xenial-inap-mtl01-0001077053.
21:09:09 mriedem Jul 31 17:18:26.237174 ubuntu-xenial-inap-mtl01-0001077052 nova-compute[2320]: DEBUG nova.virt.libvirt.driver [None req-994d9893-545f-47b9-b93e-1e21cb439db7 tempest-LiveMigrationTest-233418614 tempest-LiveMigrationTest-233418614] [instance: 7f68c430-f565-433a-8f87-27b9a00d29a0] Starting monitoring of live migration {{(pid=2320) _live_migration /opt/stack/new/nova/nova/virt/libvirt/driver.py:7555}}
21:09:15 mriedem start monitoring the live migration there ^
21:09:59 mriedem live migration complete:
21:09:59 mriedem Jul 31 17:18:27.486275 ubuntu-xenial-inap-mtl01-0001077052 nova-compute[2320]: INFO nova.compute.manager [None req-14ccce2e-8610-47b7-aba2-77f6fc468b61 tempest-LiveMigrationRemoteConsolesV26Test-255629851 tempest-LiveMigrationRemoteConsolesV26Test-255629851] [instance: 7f68c430-f565-433a-8f87-27b9a00d29a0] VM Migration completed (Lifecycle Event)
21:10:12 mriedem heh 1 second?
21:10:52 dansmith makes sense.. I doubt a cirros guest has more than a hundred meg of dirty ram
21:11:02 dansmith which is 1 second at gigE
21:11:13 mriedem i never see "VIF events received, continuing migration"
21:11:32 dansmith it's not LB
21:11:35 dansmith right?
21:11:39 mriedem oh right duh
21:11:41 prometheanfire you'll probably want to retitle the bug https://bugs.launchpad.net/nova/+bug/1786346
21:11:41 openstack Launchpad bug 1786346 in OpenStack Compute (nova) "live migrations slow" [Undecided,New]
21:11:44 prometheanfire mriedem: dansmith ^
21:11:49 dansmith prometheanfire: thanks
21:11:59 prometheanfire if you can let me know when you update the bug with details I'd appreciate it
21:12:20 dansmith I'm trying to get the revert to even pass tests and then I will
21:14:03 prometheanfire thanks
21:18:25 mriedem dansmith: ok so https://review.openstack.org/553608 should do the wait in compute now
21:19:09 dansmith mriedem: cool, updating the bug now and working on the revert in parallel, so we can make that depend on the revert to be sure we don't get the timeout message at least right?
21:19:39 mriedem well, that's why i was looking at timings,
21:20:02 mriedem because this means we'll consume the network-vif-plugged event from pre_live_migration before we call driver.live_migration which does the bw stuff,
21:20:10 mriedem the event won't come for that waiter,
21:20:27 mriedem but the guest transfer is so fast, won't we just finish the operatoin before we ever had a chance to timeout?
21:20:41 mriedem like, do i need a patch that puts a fake sleep in the driver's live migratoin metohd?
21:20:42 mriedem *method
21:21:15 dansmith even at 1MB/s?
21:21:19 mriedem could set the vif_plugging_timeout to like 1 minute, and add a 30 second sleep in the driver
21:21:28 mriedem well,
21:21:39 dansmith should go slower there, but I guess it won't take long enough
21:21:41 mriedem maybe not, but the test will timeout before the 5 minute vif_plugging_timeout i think
21:21:56 dansmith yeah okay so we'll have to force it down I guess
21:22:00 mriedem just wondering if i should set the vif_plugging_timeout to like 1 minute
21:23:15 dansmith 30sec but yeah
21:23:44 mriedem ok updated; hopefully my zuul fu is strong enough
21:23:44 jaypipes melwitt, dansmith: regarding https://review.openstack.org/#/c/540258, even if we fix the scheduler/top-level issues around server group affinity and multiple cells, that's still not going to fix the eleventh-hour on-the-compute-node checks that currently run just for affinity groups, though, right? I mean, the computes can't talk cross-cell anyway so there would be no way for those on-compute-node checks to run...
21:24:58 mriedem jaypipes: yes https://review.openstack.org/#/c/540258/8/nova/scheduler/utils.py@738
21:25:11 mriedem "Also note that we could be racing if we have multiple server create requests for the same affinity group and the scheduler decides to put them each in different cells - the late affinity check in the compute won't resolve that because the upcall check is targeted to the cell the compute is in, and won't see any other hosts for other members in other cells. Separate bug though..."
21:26:42 jaypipes mriedem: ack, ok, just wanted to verify I wasn't crazypants.

Earlier   Later