| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2019-02-28 | |||
| 14:28:40 | mdbooth | However, if all the user is trying to do is DR an instance (from the spec), going via glance isn't the most efficient solution. | |
| 14:29:22 | openstackgerrit | Jim Rollenhagen proposed openstack/nova master: Remove TypeError handling for get_info https://review.openstack.org/640043 | |
| 14:29:22 | openstackgerrit | Jim Rollenhagen proposed openstack/nova master: Fix typo in vmware get_info docstring https://review.openstack.org/640042 | |
| 14:29:52 | mdbooth | That's a point... in the DR case you'd presumably be evacuating anyway, so you still couldn't evacuate with a different root volume. | |
| 14:30:05 | sean-k-mooney | mdbooth: that depend some of the usescase in the spec impliend at least to me that the remote fail over sigt may not have direct conectivity at the hyperviosr level. e.g. a migrtion to the second sight may now actully work | |
| 14:31:15 | sean-k-mooney | *sigt -> site sight->site | |
| 14:32:03 | mdbooth | Shelve/unshelve is *extremely* space inefficient. You're basically storing a flattened snapshot at the point of shelve in the image cache forever. | |
| 14:32:29 | mdbooth | But not for BFV, I guess. | |
| 14:32:40 | sean-k-mooney | mdbooth: doesnt ath depend on your glance backend | |
| 14:33:35 | mdbooth | sean-k-mooney: No. It might depend on Nova's imagebackend, though. | |
| 14:34:23 | sean-k-mooney | ah right. i was thinking it might be effienet on ceph backed glance but may ceph imagebackend would help | |
| 14:34:40 | sean-k-mooney | e.g. for non bfv instnaces | |
| 14:35:07 | mdbooth | I don't think we're at all clever about shelve of ceph instance to ceph glance. | |
| 14:35:30 | mdbooth | And ceph glance to ceph instance is efficient but problematic. | |
| 14:38:04 | efried | gibi, bauzas: Sorry, I meant to mention it in the change set - feel free to squash. | |
| 14:41:37 | bauzas | efried: thanks | |
| 14:55:34 | artom | mriedem, hey, just to get a feel for what I should concentrate on next, is the absence of func tests for NUMA live migration a blocker? If so, I'll start on that now, otherwise, I'll advance integration testing - and can add func tests after FF? | |
| 14:56:04 | artom | (Assuming it merges, that is - but that's just my carebear optimism ;) | |
| 14:56:38 | mriedem | artom: did you see my comments about the compute service version check in the conductor task? | |
| 14:56:50 | artom | mriedem, did, and addressed in a new patch on top | |
| 14:57:00 | artom | (Figured it's legit to split them for easier reviewing) | |
| 14:57:25 | mriedem | artom: ok i didn't know how your integration testing could have even been working with that code the way it was, unless you were enabling that workaround | |
| 14:57:48 | artom | mriedem, that's... actually a good question | |
| 14:58:05 | artom | Because I sure as hell didn't set enable_numa_live_migration, and its default is False, right? | |
| 14:58:38 | mriedem | correct | |
| 14:58:50 | artom | And yet my integration tests *do* work, because they caught a logic flaw and failed when the pinning of 2 instances ended up overlapping | |
| 14:59:00 | mriedem | were you testing with that patch reverted? | |
| 14:59:03 | sean-k-mooney | there is a bug | |
| 14:59:14 | artom | mriedem, nope :/ | |
| 14:59:17 | sean-k-mooney | in the numa migration config code | |
| 14:59:35 | sean-k-mooney | https://review.openstack.org/#/c/635350 | |
| 14:59:49 | mriedem | you mean this if hypervisor_type != obj_fields.HVType.KVM: | |
| 14:59:50 | sean-k-mooney | mriedem: artom cfriesen found and fixed it | |
| 15:00:16 | sean-k-mooney | yep it will be QEMU not kvm | |
| 15:00:24 | mriedem | right so the blocker config doesn't actually do anything | |
| 15:00:30 | sean-k-mooney | right | |
| 15:00:32 | artom | I swear I see stuff like that and I can't help but think "we're all just morons" | |
| 15:00:41 | artom | Nothing against anyone, and I'm including myself in that group | |
| 15:01:08 | tacco | hi there.. i have a issue with a redeployed hypervisor.. after deployment i see a "ResourceProviderCreationFailed: Failed to create resource provider" error when trying to start nova or migrating a instance to this hypervisor https://pastebin.com/zMHUUYEW | |
| 15:01:10 | artom | At least we found and fixed it before release :) | |
| 15:01:32 | tacco | any idea? | |
| 15:01:35 | sean-k-mooney | its not merged yet. imgoing to recheck the patch. that said if we merger your code this is not needed | |
| 15:01:56 | artom | sean-k-mooney, still needed for anything but full Stein deployments | |
| 15:02:11 | artom | See https://review.openstack.org/#/c/640021/ :) | |
| 15:02:12 | sean-k-mooney | ya that is ture | |
| 15:02:16 | mriedem | artom: it's already been backported to stable/rocky | |
| 15:02:46 | artom | mriedem, the CONF workaround? *sigh* | |
| 15:03:59 | sean-k-mooney | ya so we will have to backport the fix too but sice its currently a noop it wont break anything | |
| 15:04:17 | sean-k-mooney | its annoying but thats all | |
| 15:05:06 | mriedem | artom: yes https://review.openstack.org/#/q/I217fba9138132b107e9d62895d699d238392e761 | |
| 15:05:29 | mriedem | me thinks we should probably start an rc-potential etherpad... | |
| 15:05:34 | sean-k-mooney | cfriesen: do you have time to repsin that by the way? if not ill add https://review.openstack.org/#/c/635350 to my list and ill file a bug | |
| 15:05:38 | mriedem | because there is just too much crap flying around for me to keep in my brain | |
| 15:06:51 | artom | mriedem, you mean rc blocker? | |
| 15:08:05 | mriedem | https://etherpad.openstack.org/p/nova-stein-rc-potential | |
| 15:08:08 | mriedem | melwitt: ^ | |
| 15:08:09 | sean-k-mooney | artom: not so much blocker but might warrent a second rc or shoudl be merged after ff as part of an rc | |
| 15:09:01 | artom | mriedem, so, sorry to pester, but the func tests blocker thing wasn't resolved | |
| 15:09:59 | mriedem | you want closure from me in other words? | |
| 15:10:26 | artom | mriedem, heh, yeah. A sense of priorities, o wise leader :) | |
| 15:12:44 | mriedem | well i personally think it's pretty risky to land something as complicated as live migration with numa without some functional tests, but at the same time, i guess if it turns out to have bugs then they just get fixed and we trust your integration testing downstream | |
| 15:12:56 | cfriesen | sean-k-mooney: will respin | |
| 15:13:12 | mriedem | in general i loathe landing anything related to pci/numa/sriov when we have no 3rd party integration testing for it | |
| 15:13:21 | cfriesen | we'll be doing integration testing for live migration too | |
| 15:13:35 | dansmith | mriedem: I too think it's fairly risky | |
| 15:13:41 | sean-k-mooney | cfriesen: ok it just need a bug added. | |
| 15:13:55 | sean-k-mooney | cfriesen: so it can be backported | |
| 15:14:02 | dansmith | I haven't really had time to go back through it from a holistic "how do these pieces fit together" sort of view | |
| 15:14:04 | artom | mriedem, I know :/ Which is why I want to replace the Intel CI with something that's actually maintained | |
| 15:14:17 | cfriesen | dansmith: mriedem: totally risky. but is it likely to break anything other than things that are already broken? | |
| 15:14:21 | dansmith | breaking live migration at all is a really bad thing to do | |
| 15:14:29 | dansmith | cfriesen: yes? | |
| 15:15:02 | artom | dansmith, basic live migration is tested in the gate though, right? | |
| 15:15:10 | mriedem | yes, | |
| 15:15:19 | dansmith | artom: during several partial upgrade situations? | |
| 15:15:31 | mriedem | so if worst case scenario you break live migration with numa instances, well, those were already broken anyway - and if you're using that config (once fixed) you can't even initiate it | |
| 15:15:34 | artom | No :( | |
| 15:15:40 | mriedem | dansmith: yes | |
| 15:15:44 | artom | Yes? | |
| 15:15:46 | mriedem | the live migration + grenade job | |
| 15:15:48 | dansmith | mriedem: we have one right? | |
| 15:15:50 | artom | \o/ | |
| 15:15:52 | mriedem | runs n-1 live migration back and forth | |
| 15:15:54 | mriedem | it's non-voting | |
| 15:16:13 | artom | We should *really* check its output on the NUMA LM patches then | |
| 15:16:15 | dansmith | mriedem: but not on systems that will report this stuff, even for non-numa flavors I think | |
| 15:16:24 | mriedem | https://review.openstack.org/#/c/637231/ | |
| 15:16:48 | mriedem | dansmith: correct we don't do any numa stuff in the gate | |
| 15:17:03 | sean-k-mooney | artom: i would start on the functional tests and we can add the whitebox tempest test after ff | |
| 15:17:04 | dansmith | it's not just numa-having instances, | |
| 15:17:33 | dansmith | it's that everyone's systems report numa stuff, and if any of this works on a numaless guest instance for CI but is weird on a real system... we won't have coverage of that | |
| 15:17:34 | mriedem | mrhillsman: would it be possible to use openlab for numa testing for nova? | |
| 15:17:37 | sean-k-mooney | artom: wether they are a blocker on not is seperate form if we should write them an i think we would all like to see functional tests for this eventually | |
| 15:18:41 | artom | sean-k-mooney, oh yeah, they'll be there | |
| 15:19:09 | dansmith | also, some amount of live migration of numa-having instances works today but just doesn't claim or do the right thing, right? | |
| 15:19:14 | artom | I'm wondering more "what gives others better confidence in the code to help it merge, automated integration tests that you have to run manually in your own env, or func tests?" | |
| 15:19:29 | dansmith | if we were to break that in a way that doesn't let people move their instances in an emergency, even if the numaness gets messed up, that's still a big problem I think | |
| 15:19:55 | mrhillsman | mriedem yes sir | |
| 15:20:28 | sean-k-mooney | dansmith: if the cpus are free on the destination then today everything works for pinned instance | |
| 15:20:37 | sean-k-mooney | similarly for hugepage instnace | |