| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-09-25 | |||
| 11:17:37 | gibi | jaypipes: I will do that as well as it is soo inconvenient to switch to the hu keyboard layout to type those exta dots over letters :P | |
| 11:17:37 | adrianc | stephenfin: you can use mstflint package (use mstconfig utility to enable SR-IOV), mlnxofedinstall is used to install MLNX_OFED package which includes mlxconfig utility. | |
| 11:18:23 | gibi | jaypipes: be aware that google translate only trusty for simple expressions and standalon words and it tend to fail hard on grammar :) | |
| 11:21:16 | jaypipes | gibi: :) | |
| 11:21:23 | openstackgerrit | Lee Yarwood proposed openstack/nova stable/rocky: Validate transport_url in nova-manage cell_v2 commands https://review.openstack.org/605040 | |
| 11:21:56 | openstackgerrit | Lee Yarwood proposed openstack/nova stable/rocky: Wait for network-vif-plugged on resize revert https://review.openstack.org/605041 | |
| 11:26:48 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Consumer gen: remove_provider_from_instance_allocation https://review.openstack.org/591784 | |
| 11:36:05 | sean-k-mooney | jaypipes: o/ quick question, when opening a blueprint should i set the series goal to stein or leave that to the approver when they review? | |
| 11:36:30 | sean-k-mooney | similarly for milestone target | |
| 11:41:28 | sean-k-mooney | stephenfin: gibi efried same question ^ i would like to target https://blueprints.launchpad.net/nova/+spec/libvirt-neutron-sriov-livemigration to stein and m2 but not sure if its rude too do that myself or should i leave that to the approver if approved? | |
| 11:42:46 | jaypipes | sean-k-mooney: leave that to the approver | |
| 11:43:11 | sean-k-mooney | jaypipes: cool will do thanks | |
| 11:46:46 | openstackgerrit | Matthew Booth proposed openstack/nova master: Add volume-backed evacuate test https://review.openstack.org/604397 | |
| 11:54:43 | openstackgerrit | Przemyslaw Lal proposed openstack/os-vif master: add nested DPDK VIF classes for kuryr-kubernetes https://review.openstack.org/604992 | |
| 12:15:46 | bauzas | jaypipes: efried: thanks for the super fast review of https://review.openstack.org/#/c/599208/ | |
| 12:16:00 | bauzas | I'll provide a new revision hopefully today EOB | |
| 12:16:13 | openstackgerrit | Merged openstack/nova stable/rocky: nova-status - don't count deleted compute_nodes https://review.openstack.org/604785 | |
| 12:16:15 | bauzas | that said, I need to AWOL for 30 mins (paperwork) | |
| 12:19:13 | jaypipes | bauzas: np | |
| 12:23:55 | openstackgerrit | Lee Yarwood proposed openstack/nova master: scheduler: Increase alternate count in smaller environments https://review.openstack.org/593074 | |
| 12:23:56 | openstackgerrit | Lee Yarwood proposed openstack/nova master: scheduler: rename index to num_instances https://review.openstack.org/605054 | |
| 12:29:15 | mdbooth | gibi: We jumped the gun on my notification timeout fix :) | |
| 12:29:53 | lyarwood | bauzas: ^ would you mind taking another swing at the above scheduler bugfix? | |
| 12:29:57 | mdbooth | TimeoutError isn't a thing in python2, apparently | |
| 12:30:34 | mdbooth | But the python3 functional test actually uncovered an error in another test \o/ | |
| 12:30:55 | gibi | mdbooth: interestingly I did google that exception before I +2d the patch and it seemd legit | |
| 12:31:10 | mdbooth | gibi: Should probably increase that timeout globally, tbh. In the notification tests it's 10 secs. | |
| 12:31:12 | gibi | mdbooth: then I suggest to use AssertionError as unittest asserts raises that as well | |
| 12:31:25 | mdbooth | gibi: ack. Will change it now. | |
| 12:31:32 | mdbooth | gibi: What do you think about increasing the timeout? | |
| 12:31:39 | mdbooth | http://logs.openstack.org/59/604859/2/check/nova-tox-functional-py35/1a5d5b7/job-output.txt.gz | |
| 12:31:47 | mdbooth | ^^^ is a failure in python3 functional | |
| 12:32:04 | gibi | mdbooth: if that helps with those test locally for you then I'm fine with it | |
| 12:32:22 | mdbooth | I know we hit these regularly, because I've been scratching my head about random failures in my evacuate test for a couple of weeks | |
| 12:32:33 | mdbooth | Until I found this | |
| 12:32:48 | gibi | mdbooth: ohh this is the root cause of the evac failure. good catch! | |
| 12:34:11 | mdbooth | gibi: Ok, will change to AssertionError and bump timeout to 10s | |
| 12:35:09 | mdbooth | I *almost* wrote a meta-testcase, but didn't. | |
| 12:35:19 | mdbooth | I think I'll write a small meta-testcase this time, too. | |
| 12:41:12 | openstackgerrit | Kashyap Chamarthy proposed openstack/nova master: Pick next minimum libvirt / QEMU versions for "T" release https://review.openstack.org/605060 | |
| 12:48:37 | stephenfin | adrianc: Do 'mlxconfig' and 'mstconfig' fulfil the same role, in that case? | |
| 12:51:41 | adrianc | stephenfin: Yes | |
| 13:28:15 | openstackgerrit | Eric Fried proposed openstack/nova master: doc:update virtual gpu doc https://review.openstack.org/605012 | |
| 13:32:50 | efried | mdbooth: Why is that suddenly failing in every spot?? | |
| 13:32:57 | efried | (I mean even in py3) | |
| 13:33:29 | mdbooth | efried: Which one? | |
| 13:33:36 | mdbooth | My functional test? | |
| 13:33:39 | mdbooth | for evac? | |
| 13:33:47 | efried | yeah | |
| 13:33:56 | mdbooth | Because I rebased it on top of a patch with a bug :) | |
| 13:34:09 | mdbooth | which I'm just fixing, btw | |
| 13:34:30 | efried | I'm confused. It appears to be on bottom. | |
| 13:34:42 | mdbooth | Nah, there's a new one on the bottom | |
| 13:34:51 | mdbooth | Raise error on timeout in wait_for_versioned_notifications | |
| 13:35:34 | mdbooth | I broke it this morning | |
| 13:35:47 | efried | I'm looking at https://review.openstack.org/#/c/604859/ itself. | |
| 13:36:12 | mdbooth | Yeah, there are 2 problems with that patch | |
| 13:36:18 | efried | Forget the py2 - TimeoutError ENOENT - I'm looking at the py3 failures and can't figure out why they are *all* failing now. | |
| 13:36:21 | mdbooth | First is a py2 error (novel) | |
| 13:36:38 | mdbooth | Right, that's because it's uncovering bugs in other tests | |
| 13:37:02 | efried | But even the ones that were already checking the result are now failing. | |
| 13:37:06 | mdbooth | Well, it uncovered one of those | |
| 13:37:18 | mdbooth | Wait, has it run again? I only saw a single failre | |
| 13:37:49 | efried | omg, is this going to fix the famed and fabled bug #1763181 ?? | |
| 13:37:50 | openstack | bug 1763181 in OpenStack Compute (nova) "test_parallel_evacuate_with_server_group intermittently fails" [Medium,Confirmed] https://launchpad.net/bugs/1763181 | |
| 13:37:59 | mdbooth | efried: It's probably because the gate is currently rammed | |
| 13:38:32 | efried | okay, well, I'll leave you to it. | |
| 13:38:45 | mdbooth | efried: i.e. I suspect they're actually races | |
| 13:39:05 | mdbooth | efried: I'm going to check properly, though | |
| 13:39:48 | efried | okay, never mind what I said about the existing ones failing. They're not. | |
| 13:47:32 | openstackgerrit | Matthew Booth proposed openstack/nova master: Raise error on timeout in wait_for_versioned_notifications https://review.openstack.org/604859 | |
| 13:49:14 | mdbooth | efried: Incidentally, the other possibility is that in some of these tests the notification has never been emitted | |
| 13:49:29 | mdbooth | Because that would previously have been silently ignored | |
| 13:51:14 | efried | yup | |
| 13:51:35 | mdbooth | I've just thrown that patch back into the gate with fixes, but it's going to fail again | |
| 13:51:50 | efried | mdbooth: But IMO if we want that to be okay, we should explicitly try/except+ignore the wait_for_notifications call in the test case. | |
| 13:52:02 | mdbooth | efried: ack, for sure | |
| 13:52:33 | efried | so ++ to your change, at least in principle :) | |
| 14:08:09 | mdbooth | efried gibi: In the case of nova.tests.functional.regressions.test_bug_1735407.TestParallelEvacuationWithServerGroup.test_parallel_evacuate_with_server_group the test appears to be incorrect | |
| 14:08:38 | mdbooth | It's calling evacuate on 2 instances, one of which is going to fail due to anti-affinity | |
| 14:08:57 | mdbooth | It's then asserting it got 2 instances of instance.rebuild.start | |
| 14:09:14 | mdbooth | Problem is that request validation happens before that notification is sent | |
| 14:09:26 | mdbooth | So it has never received those notifications | |
| 14:09:56 | mdbooth | However, it does sound wrong that the notification isn't emitted first | |
| 14:09:57 | gibi | mdbooth: doesn't this test want to assert that the late validate of the server groups catches the parallel evacuation? | |
| 14:10:29 | mdbooth | I believe so, yes | |
| 14:11:02 | mdbooth | However, I think it's expecting the start notification to be emitted before the failure occurs | |
| 14:11:06 | mdbooth | Which imho isn't unreasonable | |
| 14:11:15 | mdbooth | But that's not the reality | |
| 14:12:33 | mdbooth | See ComputeManager._do_rebuild_instance | |
| 14:12:43 | mdbooth | We call _validate_instance_group way at the top | |
| 14:12:58 | mdbooth | Notification is emitted below that | |
| 14:15:34 | gibi | mdbooth: ohh I see now. So we can expect one rebuild.start but the second evac will never reach that point | |
| 14:16:36 | mdbooth | gibi: ack | |
| 14:17:06 | gibi | mdbooth: I thin it is OK to change that to wait only for a single notification | |
| 14:17:38 | gibi | mdbooth: as the two self._wait_for_migration_status(server1, ['done', 'failed']) calls will make sure that the test waits for the evac to finish/fail | |
| 14:28:24 | mnaser | https://review.openstack.org/#/q/I811e84af46d678c3fdbf94ee400eabe659fc3d4e if anyone wants to continue with the backporting +2s | |
| 14:31:15 | mriedem | done | |
| 14:31:58 | mdbooth | mriedem: I messed with your script again, btw. Still waiting to see how it gets on in the gate. | |
| 14:33:59 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: consumer gen: move_allocations https://review.openstack.org/591810 | |