Earlier  
Posted Nick Remark
#openstack-nova - 2018-02-23
16:53:12 mriedem like aggregate meta, and flavor extra specs
16:53:23 mriedem i realize it's use though
16:54:06 cfriesen dansmith: it'd be possible to prohibit changing the value on an existing group to something that would result in the current spread being invalid....alternately you could just allow that and document that it'll only affect the *next* scheduling decision.
16:54:39 mriedem cfriesen: to determine if the new requested value would invalidate things would mean running through the scheduler all over again
16:54:44 mriedem and you could still get it wrong
16:54:53 cfriesen mriedem: there has been that recurring spec for allowing instances to be added/subtracted from a group
16:54:54 mriedem which is why we have the late affinity check on the compute hosts
16:55:12 mriedem cfriesen: i remember powervc pushing it back in kilo but that's been abandoned for a long time
16:55:39 cfriesen arguably we've got races all over with instance group affinity, so what's one more. :)
16:55:41 dansmith cfriesen: yeah then the group is in violation of policy with no way of getting it out, unless you do your own shuffling
16:56:07 dansmith mriedem: and yeah I'd rather a real attribute, but I also think we're just going to end up with unlimited attributes for other things like this
16:56:15 dansmith mriedem: so I dunno.. it's a can of worms
16:57:53 mriedem idk, if this is the first time somtehing like this has come up in the last what 5 years?
16:58:23 cfriesen for what it's worth, internally we added a "best-effort" flag to the "hard" affinity/antiaffinity policies to allow us to migrate instances off a compute node for maintenance
16:58:23 mriedem it would get annoying if, over time, we had a bunch of attributes on groups that only applied to specific policies
16:58:37 mriedem cfriesen: so soft affinity?
16:58:56 cfriesen no, you can turn the best-effort flag on and off dynamically
16:59:17 cfriesen so you'd normally run with it strict, but if you need to take down a compute node you can set best-effort, move everything, then turn it back off
16:59:41 cfriesen otherwise if you've got hard-affinity you can't migrate any of them
17:00:04 mriedem which is why people want a force flag on evacuate, live migrate (and now cold migrate)
17:00:06 leakypipes mriedem: and the late affinity check is (the only?) remaining upcall from a cell to API, no?
17:00:15 mriedem leakypipes: hells no
17:00:32 leakypipes mriedem: it's not an upcall?
17:00:35 mriedem https://docs.openstack.org/nova/latest/user/cellsv2-layout.html#operations-requiring-upcalls
17:00:38 mriedem it is an upcall
17:00:42 mriedem but it's not the only one
17:00:44 leakypipes oh, not the only reminaing..
17:00:46 mriedem we've got aggregates too
17:01:04 mriedem so aggregates and affinity are the remaining upcall issues
17:01:21 mriedem but right now there are 2 each
17:01:24 mriedem so 4 upcall issues
17:01:33 leakypipes mriedem: then you'll LOVE my aggregate affinity spec! :) now with moar AFFINITY and moar AGGREGATES!
17:01:48 mriedem leakypipes: i already said 'upcall upcall upcall' on that spec several times :)
17:01:54 leakypipes I know :)
17:02:20 mriedem i think i also hedged with something like, 'but we already do this in a few other places so people already have to rely on it, so maybe another log on the fire doesn't kill us'
17:02:55 cfriesen On a totally different topic...has anyone ever heard of nova allocating duplicate network interfaces? (So the user boots while asking for 2 network interfaces, and nova allocates two ports on each network.)
17:03:08 mriedem yes
17:03:10 mriedem that's old news
17:03:25 mriedem tempest has a test for it also i think
17:03:31 mriedem been around since juno?
17:05:28 cfriesen I mean nova is allocating twice as many as were asked for.
17:05:47 mriedem double your pleasure
17:05:49 mriedem idk, going to lunch
17:14:43 efried Greetings from JFK airport
17:27:30 mnaser ok
17:27:33 mnaser im convinced grenade is broken
17:27:43 mnaser for stable/pike
17:28:05 efried Isn't that mnaser guy known for being johnny-on-the-spot for grenade fixes?
17:28:16 mnaser only when i have to :(
17:28:29 mnaser Host 'ubuntu-xenial-rax-dfw-0002683360' is not mapped to any cell
17:28:36 mnaser we keep getting this in multinode
17:29:49 efried It would seem odd that cell discovery isn't being run. Like, nothing would ever work.
17:30:00 efried And that's pretty much the only thing I know about cells.
17:30:40 mnaser efried: indeed seems to be the case. looks like there was a change to add 'CELLSV2_SETUP=singleconductor' in there, so not sure if that might have affected it
17:31:13 efried mnaser: You're going to have something of a hard time finding a core today, but I'll +1 your fix :)
17:31:31 mnaser efried: looks like there's only 3 cores for grenade too..
17:32:49 efried Looks like qa-release is included by inheritance.
17:33:02 efried So seven
17:35:12 mnaser ok looks like this runs => nova-manage cell_v2 simple_cell_setup --transport-url rabbit://stackrabbit:secretrabbit@10.209.130.218:5672/
17:35:39 mnaser but discover_hosts is never called
17:36:22 efried And simple_cell_setup doesn't run discovery itself?
17:36:52 efried I remember having to fix this around pike timeframe.
17:37:02 mnaser efried: going through the code it looks like it does call _map_cell_and_hosts()
17:47:51 openstack Launchpad bug 1708039 in devstack "gate-grenade-dsvm-neutron-multinode-ubuntu-xenial fails with "No host-to-cell mapping found for selected host"" [Medium,Fix released] - Assigned to Sean Dague (sdague)
17:47:51 mnaser sigh https://bugs.launchpad.net/grenade/+bug/1708039 looks like it was 'supposed' to be fixed
17:49:09 mnaser it looks like it regressed
17:49:12 mnaser and its all stable/pike hits
17:54:40 mnaser "Didn't find service registered by hostname after 60 seconds" .. found it
17:54:49 mnaser it's actually listed but the bash for some reason doesnt find it
17:55:51 efried leakypipes: You around today?
17:55:58 efried Hoho, it's Friday
17:56:14 andreaf mriedem hey I'm setting up a zuul-v3 multinode job, and everything works fine apart from nova that gives me "Host is not mapped to any cell" http://logs.openstack.org/24/545724/9/check/tempest-multinode-full/1bbec81/ara/result/525a60bd-ac22-4fc4-9db7-e61fce8ac1f5/
17:56:45 andreaf mriedem: I compared configs in localrc and nova and I don't see anything obvious - do you have any idea about what this could be?
17:56:52 leakypipes fried_rice_jfk: ues
17:56:54 leakypipes yes
17:57:06 fried_rice_jfk leakypipes: It looks like alex_xu may be right. I wrote a gabbit for it.
17:57:49 mnaser andreaf: im actually looking into this right ow
17:57:55 mnaser im seeing this issue in stable/pike
17:58:38 andreaf mnaser oh ok at least it's not just me :P
17:59:00 mnaser andreaf: i'm seeing the devstack start waiting for compute to go up at "2018-02-23 04:24:33.575" (with a 60s timeout) and the compute record get created at "2018-02-23 04:25:44.182"
17:59:31 mnaser with a 60 second timeout, it means that devstack gives up at 04:25:33, but the compute record actually gets created 11 seconds later
18:01:29 fried_rice_jfk leakypipes: Trying to figure out why. Is the order in which I construct my query supposed to not matter?
18:03:26 fried_rice_jfk leakypipes: Maybe query.select_from() overwrites any previous .select_from()?
18:04:16 mnaser so it takes 42 seconds to go from "Connecting to libvirt: qemu:///system _get_new_connection /opt/stack/old/nova/nova/virt/libvirt/host.py:366" => "Registering for lifecycle events"
18:04:38 mnaser and those 42 seconds are enough for devstack to timeout waiting for compute
18:05:42 mnaser so somehow "wrapped_conn = self._connect(self._uri, self._read_only)" takes 42 seconds here https://github.com/openstack/nova/blob/stable/pike/nova/virt/libvirt/host.py#L368
18:07:52 mnaser libvirtd starts at "2018-02-23 04:23:59.741" from devstack
18:07:57 fried_rice_jfk leakypipes: Gah, my apologies; I was misreading my test failure. It's fine, it's cumulative as we expected.
18:08:42 openstackgerrit Eric Fried proposed openstack/nova master: rp: GET /resource_providers?required= https://review.openstack.org/546837
18:08:52 leakypipes fried_rice_jfk: k
18:08:54 fried_rice_jfk leakypipes, alex_xu: Added a gabbit to prove ANDness with resources ^
18:18:25 mnaser who's the person to bug for libvirt related questions
18:19:01 mnaser i'm seeing 1116 "device-list-properties" on service start
18:22:05 figleaf fried_rice_jfk: looks good!
18:22:22 fried_rice_jfk Thanks figleaf
18:25:06 openstackgerrit Eric Berglund proposed openstack/nova master: PowerVM Driver: vSCSI volume driver https://review.openstack.org/526094
19:17:12 mnaser o/
19:20:33 mriedem andreaf: mnaser: maybe on a slow test node,
19:20:58 mriedem devstack times out waiting for the compute host to get mapped to the cell, which won't happen until after the compute node record gets auto-created when the nova-compute service starts up
19:21:29 mnaser mriedem: exactly, but libvirt takes 42 seconds to start up and seems to be doing a ton of commands before starting up

Earlier   Later