Earlier  
Posted Nick Remark
#openstack-nova - 2018-10-03
15:26:53 sean-k-mooney then you could decide if it shoudl stay in error or active state seperately without haveing more bugs if you left it in active
15:27:39 openstackgerrit Claudiu Belu proposed openstack/nova master: hyper-v: autospec classes before they are instantiated https://review.openstack.org/342211
15:28:33 sean-k-mooney mriedem: or to put that another way i think there are two issues in the other case. 1 the host is not upstaed when the vm is moved in some cases and 2 what to do when cleanup fails post migration
15:28:55 mriedem if post live migration set the instance.host to the correct host on which it's running then yeah my concern about the user rebooting it and now having the same guest on different hosts is less of an issue
15:32:08 sean-k-mooney nicolasbock: if you are still around the clip notes version of that conversation is you might want to consider locking the instance untill you have reparied the db to prevent any lifecylce envents
15:33:07 mdbooth mriedem: IIRC my thought at the time was that we should completely update the instance record for the destination immediately after we switch it, then run source cleanup.
15:33:36 mdbooth So if we fail for whatever reason we've recorded that we're running on the dest.
15:37:46 mriedem s10: i think the unit test is going to fail in that patch, see comments for why
15:37:51 mdbooth Incidentally, to reiterate something I said earlier, when this patch landed late last week it made the test_parallel_evacuate_with_server_group about 20 times more likely to occur: https://review.openstack.org/#/c/604859/
15:38:08 mdbooth That test is now failing around 50% for me.
15:38:14 mriedem we can skip the test for now
15:38:20 mriedem while the fix is being reviewed
15:38:44 mdbooth mriedem: ack. Given ^^^ it seems that the test has never been good.
15:40:52 openstackgerrit Rodolfo Alonso Hernandez proposed openstack/os-vif master: Add native implementation OVSDB API https://review.openstack.org/482226
15:43:09 openstackgerrit Matt Riedemann proposed openstack/nova master: Skip test_parallel_evacuate_with_server_group until fixed https://review.openstack.org/607620
15:43:19 mriedem dansmith: efried: mdbooth: ^ gives time to be comfortable with the fix
15:44:21 mdbooth mriedem: ack.
15:44:22 efried mriedem: What will the criteria for comfort be?
15:45:05 efried mriedem: https://review.openstack.org/#/c/605436/ has three +1s and a +2
15:45:33 mriedem i guess that's up to whoever +Ws it
15:46:12 efried mriedem: Well, I'm comfortable that it fixes the problem. But don't feel confident enough in the actual code change to +W. I would think someone like.... mriedem would be able to have that confidence.
15:47:47 mriedem i'm not very confident in anything atm
15:48:28 dansmith I think that change needs a lot of inspection
15:48:35 dansmith which I can't do right this moment
15:59:45 mriedem who enjoys a good UnboundLocalError? https://github.com/openstack/nova/blob/237ced4737a28728408eb30c3d20c6b2c13b4a8d/nova/network/neutronv2/api.py#L1429
16:02:54 mriedem oh i guess it's not unbound, it's a module import...
16:09:43 mriedem so uh, if we hit ^ shouldn't we fail the build?
16:10:08 mriedem clearly the user isn't going to get the sriov port attached to the guest that they requested
16:11:12 openstackgerrit Jay Pipes proposed openstack/nova stable/ocata: Re-use existing ComputeNode on ironic rebalance https://review.openstack.org/607626
16:12:10 openstackgerrit Matt Riedemann proposed openstack/nova master: Fix logging parameter in _populate_pci_mac_address https://review.openstack.org/607628
16:12:54 mriedem sean-k-mooney: you might be interested in https://bugs.launchpad.net/nova/+bug/1795064
16:12:54 openstack Launchpad bug 1795064 in OpenStack Compute (nova) "SR-IOV error IndexError: pop from empty list" [Undecided,New]
16:13:06 mriedem something something sriov and kernel versions
16:13:26 sean-k-mooney looking
16:13:47 jaypipes mriedem: https://review.openstack.org/#/c/607626/ is that stable/ocata backport for the duplicate hypervisor_hostname thingie
16:13:56 jaypipes mriedem: thx for your help earlier.
16:15:11 mriedem jaypipes: so you cherry-picked that from master?
16:15:16 mriedem https://review.openstack.org/#/c/508555/
16:15:39 mriedem or did you cherry pick from the pike backport but forget the -x option on the cherry-pick command?
16:15:44 spatel_ Hi folks
16:15:50 jaypipes mriedem: no, I cherry-picked the SHA1 from the stable/pike patch
16:16:06 jaypipes mriedem: oh, sorry, I don't know about -x :(
16:16:28 spatel_ I am having issue with SR-IOV with shared PCI device between numa and reading this blueprint : https://blueprints.launchpad.net/nova/+spec/share-pci-between-numa-nodes
16:17:25 sean-k-mooney spatel_: this bug https://bugs.launchpad.net/nova/+bug/1795064? or another?
16:17:25 openstack Launchpad bug 1795064 in OpenStack Compute (nova) "SR-IOV error IndexError: pop from empty list" [Undecided,New]
16:17:25 spatel_ I have set hw:pci_numa_affinity_policy='preferred' in flavor but still its not allowing me to run instance on NUMA-1
16:17:58 spatel_ sean-k-mooney: that problem got resolved by downgrading kernel to 3.x
16:18:39 sean-k-mooney spatel_: i think you issue with 4.18 was that you did not have a netdev associated with the vf
16:19:04 spatel_ Is that configuration issue or BUG?
16:19:46 sean-k-mooney spatel_: i would say config issue. i would guess the default options for the gernel module cahnged and or you are using a different driver by default
16:19:47 jaypipes mriedem: apologies. how can I fix appropriately? do I need to re-do the git cherry-pick with -x? or can/should I just edit the commit message with seomthing?
16:20:27 sean-k-mooney spatel_: for example if the device was bound to vfio_pci instead fo the broadcom driver then it would existit in lspci but not have a netdev
16:20:59 openstackgerrit Artom Lifshitz proposed openstack/nova master: WIP: Libvirt live migration: update NUMA XML for dest https://review.openstack.org/575179
16:20:59 openstackgerrit Artom Lifshitz proposed openstack/nova master: Service version check for NUMA live migration https://review.openstack.org/566723
16:21:01 spatel_ hmmm!
16:21:02 sean-k-mooney spatel_: do you whitelist devices using the devname option?
16:21:04 mriedem jaypipes: see my other comment in the ocata backport about documenting the conflicts?
16:21:18 spatel_ sean-k-mooney: yes i am using devname option to specify my interface
16:21:29 spatel_ pci_passthrough_whitelist = "{ "physical_network":"vlan", "devname":"eno2" }"
16:21:36 sean-k-mooney spatel_: in general i advise against that for this exact reason
16:21:52 sean-k-mooney spatel_: if you use the pci adddress instead 4.18 would likely be fine
16:22:14 spatel_ pci address ?
16:22:27 sean-k-mooney spatel_: the whitelist can have 3 modes of whitelisting
16:22:35 spatel_ you mean vendor_id or product_id ?
16:23:09 sean-k-mooney spatel_: you can used devname, (vendor_id and product_id) or you can pass a pci address
16:23:24 spatel_ sean-k-mooney: i will give it a try and report back to BUG
16:23:41 spatel_ sean-k-mooney: currently i am dealing with this issue :( https://bugs.launchpad.net/nova/+bug/1795920
16:23:41 openstack Launchpad bug 1795920 in OpenStack Compute (nova) "SR-IOV shared PCI numa not working " [Undecided,New]
16:23:58 spatel_ Do you know what wrong i am doing here
16:24:14 spatel_ I have 2 NUMA node and running SR-IOV with shared PCI
16:24:18 sean-k-mooney https://docs.openstack.org/mitaka/networking-guide/config-sriov.html has the doc
16:24:23 sean-k-mooney am let me look
16:24:33 spatel_ I am only able to use one side of NUMA
16:25:27 sean-k-mooney by default unless you set a pci numa affinity policy in the flavor or image we require strict numa afinity
16:25:32 spatel_ Its not allowing me to launch SR-IOV instance on NUMA-2 ( because PCI is attach to NUMA-1 )
16:25:52 spatel_ All i did is hw:pci_numa_affinity_policy=preferred in flavor
16:25:57 spatel_ what else i need to do?
16:26:39 sean-k-mooney spatel_: let me check. i taught that was enough but you might also need to set the policy in the whitelist
16:26:40 spatel_ I do have aggregate_instance_extra_specs:pinned='true', hw:cpu_policy='dedicated' in flavor
16:27:10 spatel_ I think document isn't clear in blueprint so i am totally confused :(
16:28:06 spatel_ If i remove "aggregate_instance_extra_specs:pinned='true', hw:cpu_policy='dedicated'" from flavor then i am able to launch instance anywhere in NUMA with SR-IOV support
16:29:03 sean-k-mooney aggregate_instance_extra_specs:pinned='true' is not a standard thing
16:30:03 spatel_ hmmm! i didn't get it
16:30:20 sean-k-mooney spatel_: you should not need to set anything in the aggragte to use the pci policies
16:30:29 openstackgerrit Jay Pipes proposed openstack/nova stable/ocata: Re-use existing ComputeNode on ironic rebalance https://review.openstack.org/607626
16:30:46 spatel_ oh! so you are saying i should remove aggregate_instance_extra_specs:pinned
16:30:55 jaypipes mriedem: k, hopefully correct now.
16:31:02 jaypipes thx for the help again.
16:31:20 sean-k-mooney spatel_: yes
16:31:31 spatel_ lets say if i remove "aggragte" then does my vCPU get Pinned ?
16:33:57 spatel_ removing aggrate and going to launch instance
16:36:25 spatel_ sean-k-mooney: didn't work error 'No valid host was found. There are not enough hosts available'
16:36:49 spatel_ look like something is missing..
16:38:38 sean-k-mooney spatel_: so jsut to confirm you dont have any aggragte metadata set and have hw:pci_numa_affinity_policy=preferred set
16:38:48 sean-k-mooney in the flavor
16:39:15 spatel_ This is what i have currently in flavor -> properties | hw:cpu_policy='dedicated', hw:numa_nodes='2', hw:pci_numa_affinity_policy='preferred'
16:39:57 spatel_ If i remove all 3 option then i am successfully able to launch instance
16:39:57 stephenfin spatel_, sean-k-mooney: We didn't implement it with a flavor extra spec in the end
16:40:13 sean-k-mooney so this has be set in the pci whitelist then
16:40:17 stephenfin spatel_: Yes

Earlier   Later