Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-22
16:48:35 mriedem yeah i know
16:48:37 superdan okay
16:48:37 superdan for each instance in a list
16:48:48 mriedem just remembered it being a thing once
16:51:06 leakypipes looking now.
16:52:48 superdan mriedem: leakypipes: note that right now we're getting instance faults either one by one, or later from the api if it does a fill_faults on the InstanceList
16:53:04 superdan this is really just doing the latter from the lower layer if it was requested
16:53:27 superdan fill_faults from the api the way it is today won't work because it doesn't know about things in different cells
16:53:50 superdan which is a bug in pike I guess, but fixing it would not be super easy
16:54:04 mriedem hmm, probably worth at least reporting that so we know about it
16:54:13 mriedem could also doc as a known issue if nothing else
16:54:16 mriedem *always
16:54:27 mriedem shite, lunch tie
16:54:28 mriedem *time
16:54:35 superdan lemme see if I can write a test to poke that
17:00:38 figleaf superdan: the claimed hosts will still be in the list unless a filter tosses them. Since we aren't using resource filters, it's not likely that it would be removed from the list of hosts, but the sorting for alternates will certainly change
17:01:42 superdan figleaf: okay, but would it remove hosts that no longer fit after the primary consume operation was made?
17:02:11 superdan figleaf: meaning if we choose an almost-full host as a primary for one instance and consume_from_instance it or whatever, is it no longer a candidate for the last loop at the bottom to look for alternates?
17:02:46 figleaf superdan: depends on the weighers used, right? E.g., pack-vs-spread
17:03:07 openstackgerrit Kashyap Chamarthy proposed openstack/nova-specs master: [WIP] Add ability for OVMF Secure Boot https://review.openstack.org/506720
17:03:44 figleaf superdan: it won't get removed, since resource constraints are handled in the placement call
17:03:47 superdan figleaf: well I meant if we consumed some resource, but I guess we don't have things like RamFilter in there by default anymore, so this just randomizes the result (if configured) and could hit all the same ones
17:03:49 superdan yeah
17:04:17 superdan hrm. seems like we're asking for trouble there
17:04:53 figleaf superdan: well, none of the claimed hosts will be added as alternates
17:05:15 leakypipes superdan: did you catch my request on the previous revision of your drop_allocation_for_move() patch about not using the internal RT compute_nodes dict data?
17:05:16 superdan figleaf: that's what I was just asking.. what prevents that?
17:05:28 superdan leakypipes: apparently not
17:05:51 figleaf superdan: Line 285: if host.cell_uuid == cell_uuid and host not in claimed_hosts:
17:06:03 superdan figleaf: heh, oh.. THAT
17:06:24 superdan I just saw the cell check and short-circuited the rest of the line
17:06:50 figleaf superdan: I can switch the order of that line if it helps :)
17:06:56 superdan no
17:09:20 leakypipes superdan: added some more info to the latest rev
17:09:30 superdan leakypipes: okay
17:09:31 superdan thanks
17:15:31 jose-phillips d
17:21:41 superdan mriedem: oh, no it'll work because if we fail to fill, we go one by one and grab the fault, targeted to the cell
17:21:48 superdan which is terribad, but it'll work
17:25:57 openstackgerrit Sean Dague proposed openstack/nova master: Change livesnapshot to true by default https://review.openstack.org/454323
17:33:14 figleaf leakypipes: it's not just tests that use dicts from _schedule_instances: https://github.com/openstack/nova/blob/master/nova/conductor/manager.py#L572-L573
17:34:03 leakypipes figleaf: removing that code block successfully identifies the culprits though, no? :)
17:35:05 figleaf leakypipes: yeah, but it means changing it here, and then in a later patch, changing it again to use Selection objects.
17:35:09 openstackgerrit Dan Smith proposed openstack/nova master: Make allocation cleanup honor new by-migration rules https://review.openstack.org/498948
17:35:09 openstackgerrit Dan Smith proposed openstack/nova master: Move allocation manipulation out of drop_move_claim() https://review.openstack.org/498947
17:35:10 openstackgerrit Dan Smith proposed openstack/nova master: Revert allocations by migration uuid https://review.openstack.org/498949
17:35:10 openstackgerrit Dan Smith proposed openstack/nova master: Pre-create migration object https://review.openstack.org/498950
17:35:11 openstackgerrit Dan Smith proposed openstack/nova master: WIP: Make migration uuid hold allocations for migrating instances https://review.openstack.org/506420
17:35:11 openstackgerrit Dan Smith proposed openstack/nova master: Refactor resource tracker to account for migration allocations https://review.openstack.org/506419
17:35:12 openstackgerrit Dan Smith proposed openstack/nova master: Add get_node_uuid() helper to ResourceTracker https://review.openstack.org/506730
17:35:19 figleaf leakypipes: So what if we leave this, knowing that it will be removed in a later patch?
17:35:29 figleaf this == host state dicts
17:35:57 leakypipes figleaf: that's cool with me. just make a note to address in a later patch is fine.
17:36:15 openstackgerrit Merged openstack/nova master: Fix 500 if list servers called with empty regex pattern https://review.openstack.org/506585
17:36:20 figleaf leakypipes: okie dokie
17:38:11 superdan figleaf: leakypipes: best way to do that is throw a patch up with the removal at the end
17:38:18 superdan clearly broken, but a tombstone reminder :)
17:38:28 leakypipes superdan: yep
17:38:44 superdan otherwise I doubt we'll come back to it
17:43:27 figleaf superdan: not sure that's necessary, since we'll be changing the return value from a host (dict or HostState) to a Selection object. All the code that touches select_destinations() will be affected
17:45:08 superdan figleaf: okay I guess I thought this was in the middle before the point at which we'd be forming out Selection objects
17:45:41 superdan I misread your "knowing it will be removed in a later patch" as "I'll clean it up later"
17:46:23 sdague mriedem: ok, the qemu patch, who else should take a look at it - https://review.openstack.org/#/c/505673/ ?
17:46:29 figleaf superdan: this was a small change to accomodate a list of objects instead of a single, because alternates
17:46:41 sdague as it would be good to unblock new qemu
17:46:43 figleaf later will be a list of Selection
17:57:25 superdan leakypipes: just a +W needed on a test add: https://review.openstack.org/#/c/505392/7
17:57:58 openstackgerrit Jackie Truong proposed openstack/nova master: Add trusted_certs to instance_extra https://review.openstack.org/457711
18:22:07 sdague cburgess: responded on https://review.openstack.org/#/c/505673
18:22:19 leakypipes superdan: done
18:23:30 openstackgerrit Jay Pipes proposed openstack/nova master: placement: integrate ProviderTree to report client https://review.openstack.org/415921
18:23:30 openstackgerrit Jay Pipes proposed openstack/nova master: placement: set/check if inventory change in tree https://review.openstack.org/470575
18:23:31 openstackgerrit Jay Pipes proposed openstack/nova master: placement: allow filter providers in tree https://review.openstack.org/377215
18:23:31 openstackgerrit Jay Pipes proposed openstack/nova master: placement: add nested resource providers https://review.openstack.org/377138
18:23:32 openstackgerrit Jay Pipes proposed openstack/nova master: placement: update client to set parent provider https://review.openstack.org/385693
18:23:32 openstackgerrit Jay Pipes proposed openstack/nova master: placement: adds REST API for nested providers https://review.openstack.org/384807
18:36:46 cburgess sdague Re-responded
18:37:22 sdague cburgess: you all really mix / match qemu runtime and tooling versions?
18:37:30 sdague Because that kind of seems dangerous
18:37:31 cburgess In the past yes.
18:37:58 cburgess sdague Used to be required to support qemu upgrades and live migration cross versions. In theory we now claim its simply Red Hat's problem.
18:38:26 cburgess sdague I'm not saying its even a reasonable thing to do anymore. I'm simply pointing out that we are assuming something there that might not be obvious to some folks.
18:38:38 sdague cburgess: what was the sequence of changing things you would do there?
18:39:31 cburgess sdague Test the version directly.. not just ask libvirt the emulate version (assuming there is some kind of --version arguement to qemu-img).
18:40:03 sdague cburgess: there is, but then you are text parsing vs. getting a binary number, which is a lot less robust
18:40:11 cburgess sdague Maybe we should need a comment above that block of code that outlines our assumption.
18:40:37 sdague cburgess: I don't mean for this fix, I mean, lets pretend you had a 2.8 based environment, and you wanted to live migrate to a 2.10 one
18:40:48 sdague what was the series of steps metacloud would do to dot hat
18:40:55 cburgess sdague Oh you mean how did we handle this in the past?
18:40:58 sdague yes
18:41:28 openstackgerrit Merged openstack/nova master: Add fault-filling into instance_get_all_by_filters_sort() https://review.openstack.org/505391
18:43:50 openstackgerrit Merged openstack/nova master: Add a regression test for bug 1718455 https://review.openstack.org/506092
18:43:51 openstack bug 1718455 in OpenStack Compute (nova) "[pike] Nova host disable and Live Migrate all instances fail." [Medium,In progress] https://launchpad.net/bugs/1718455 - Assigned to Matt Riedemann (mriedem)
18:44:34 cburgess sdague QEMU was installed in /opt/qemu/version_numer, We had a wrapper script that lived at /usr/sbin that libvirt would find. When libvirt propped it, it would report its latest version. When it was used to launch a VM it would look at various CLI flags to determine the original version that was used to launch that VM (in the case of an incoming live migration) and launch the original version of the emulator to ensure compat.
18:45:57 sdague cburgess: and qemu-img was at what version?
18:46:23 sdague because that's in the $PATH so you do only get one there
18:46:42 cburgess sdague qemu-img was always the latest. along with the proped version from libvirt. In our case this assumption is fine. I'm not saying we (Metacloud) need this. I'm just pointing out that there is an implied assumption here.
18:46:43 sdague assuming you had 2.6, 2.8, 2.10 installed
18:46:55 sdague qemu-img was 2.10
18:47:03 cburgess Correct

Earlier   Later