| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-02-23 | |||
| 16:51:49 | mriedem | let me get you frank thomas' number | |
| 16:52:00 | mriedem | actually, jimmy johnson sells them too and he already lives in FL | |
| 16:52:32 | cfriesen | dansmith: we originally had metadata in instance groups, but it got pulled out due to not really having any users | |
| 16:52:32 | mriedem | sorry i got distracted, was hastily throwing together a PBC... | |
| 16:52:54 | mriedem | cfriesen: yeah that was linked into the spec and was something i didn't even know existed | |
| 16:53:04 | mriedem | metadata in general makes our lives terrible | |
| 16:53:12 | mriedem | like aggregate meta, and flavor extra specs | |
| 16:53:23 | mriedem | i realize it's use though | |
| 16:54:06 | cfriesen | dansmith: it'd be possible to prohibit changing the value on an existing group to something that would result in the current spread being invalid....alternately you could just allow that and document that it'll only affect the *next* scheduling decision. | |
| 16:54:39 | mriedem | cfriesen: to determine if the new requested value would invalidate things would mean running through the scheduler all over again | |
| 16:54:44 | mriedem | and you could still get it wrong | |
| 16:54:53 | cfriesen | mriedem: there has been that recurring spec for allowing instances to be added/subtracted from a group | |
| 16:54:54 | mriedem | which is why we have the late affinity check on the compute hosts | |
| 16:55:12 | mriedem | cfriesen: i remember powervc pushing it back in kilo but that's been abandoned for a long time | |
| 16:55:39 | cfriesen | arguably we've got races all over with instance group affinity, so what's one more. :) | |
| 16:55:41 | dansmith | cfriesen: yeah then the group is in violation of policy with no way of getting it out, unless you do your own shuffling | |
| 16:56:07 | dansmith | mriedem: and yeah I'd rather a real attribute, but I also think we're just going to end up with unlimited attributes for other things like this | |
| 16:56:15 | dansmith | mriedem: so I dunno.. it's a can of worms | |
| 16:57:53 | mriedem | idk, if this is the first time somtehing like this has come up in the last what 5 years? | |
| 16:58:23 | cfriesen | for what it's worth, internally we added a "best-effort" flag to the "hard" affinity/antiaffinity policies to allow us to migrate instances off a compute node for maintenance | |
| 16:58:23 | mriedem | it would get annoying if, over time, we had a bunch of attributes on groups that only applied to specific policies | |
| 16:58:37 | mriedem | cfriesen: so soft affinity? | |
| 16:58:56 | cfriesen | no, you can turn the best-effort flag on and off dynamically | |
| 16:59:17 | cfriesen | so you'd normally run with it strict, but if you need to take down a compute node you can set best-effort, move everything, then turn it back off | |
| 16:59:41 | cfriesen | otherwise if you've got hard-affinity you can't migrate any of them | |
| 17:00:04 | mriedem | which is why people want a force flag on evacuate, live migrate (and now cold migrate) | |
| 17:00:06 | leakypipes | mriedem: and the late affinity check is (the only?) remaining upcall from a cell to API, no? | |
| 17:00:15 | mriedem | leakypipes: hells no | |
| 17:00:32 | leakypipes | mriedem: it's not an upcall? | |
| 17:00:35 | mriedem | https://docs.openstack.org/nova/latest/user/cellsv2-layout.html#operations-requiring-upcalls | |
| 17:00:38 | mriedem | it is an upcall | |
| 17:00:42 | mriedem | but it's not the only one | |
| 17:00:44 | leakypipes | oh, not the only reminaing.. | |
| 17:00:46 | mriedem | we've got aggregates too | |
| 17:01:04 | mriedem | so aggregates and affinity are the remaining upcall issues | |
| 17:01:21 | mriedem | but right now there are 2 each | |
| 17:01:24 | mriedem | so 4 upcall issues | |
| 17:01:33 | leakypipes | mriedem: then you'll LOVE my aggregate affinity spec! :) now with moar AFFINITY and moar AGGREGATES! | |
| 17:01:48 | mriedem | leakypipes: i already said 'upcall upcall upcall' on that spec several times :) | |
| 17:01:54 | leakypipes | I know :) | |
| 17:02:20 | mriedem | i think i also hedged with something like, 'but we already do this in a few other places so people already have to rely on it, so maybe another log on the fire doesn't kill us' | |
| 17:02:55 | cfriesen | On a totally different topic...has anyone ever heard of nova allocating duplicate network interfaces? (So the user boots while asking for 2 network interfaces, and nova allocates two ports on each network.) | |
| 17:03:08 | mriedem | yes | |
| 17:03:10 | mriedem | that's old news | |
| 17:03:25 | mriedem | tempest has a test for it also i think | |
| 17:03:31 | mriedem | been around since juno? | |
| 17:05:28 | cfriesen | I mean nova is allocating twice as many as were asked for. | |
| 17:05:47 | mriedem | double your pleasure | |
| 17:05:49 | mriedem | idk, going to lunch | |
| 17:14:43 | efried | Greetings from JFK airport | |
| 17:27:30 | mnaser | ok | |
| 17:27:33 | mnaser | im convinced grenade is broken | |
| 17:27:43 | mnaser | for stable/pike | |
| 17:28:05 | efried | Isn't that mnaser guy known for being johnny-on-the-spot for grenade fixes? | |
| 17:28:16 | mnaser | only when i have to :( | |
| 17:28:29 | mnaser | Host 'ubuntu-xenial-rax-dfw-0002683360' is not mapped to any cell | |
| 17:28:36 | mnaser | we keep getting this in multinode | |
| 17:29:49 | efried | It would seem odd that cell discovery isn't being run. Like, nothing would ever work. | |
| 17:30:00 | efried | And that's pretty much the only thing I know about cells. | |
| 17:30:40 | mnaser | efried: indeed seems to be the case. looks like there was a change to add 'CELLSV2_SETUP=singleconductor' in there, so not sure if that might have affected it | |
| 17:31:13 | efried | mnaser: You're going to have something of a hard time finding a core today, but I'll +1 your fix :) | |
| 17:31:31 | mnaser | efried: looks like there's only 3 cores for grenade too.. | |
| 17:32:49 | efried | Looks like qa-release is included by inheritance. | |
| 17:33:02 | efried | So seven | |
| 17:35:12 | mnaser | ok looks like this runs => nova-manage cell_v2 simple_cell_setup --transport-url rabbit://stackrabbit:secretrabbit@10.209.130.218:5672/ | |
| 17:35:39 | mnaser | but discover_hosts is never called | |
| 17:36:22 | efried | And simple_cell_setup doesn't run discovery itself? | |
| 17:36:52 | efried | I remember having to fix this around pike timeframe. | |
| 17:37:02 | mnaser | efried: going through the code it looks like it does call _map_cell_and_hosts() | |
| 17:47:51 | openstack | Launchpad bug 1708039 in devstack "gate-grenade-dsvm-neutron-multinode-ubuntu-xenial fails with "No host-to-cell mapping found for selected host"" [Medium,Fix released] - Assigned to Sean Dague (sdague) | |
| 17:47:51 | mnaser | sigh https://bugs.launchpad.net/grenade/+bug/1708039 looks like it was 'supposed' to be fixed | |
| 17:49:09 | mnaser | it looks like it regressed | |
| 17:49:12 | mnaser | and its all stable/pike hits | |
| 17:54:40 | mnaser | "Didn't find service registered by hostname after 60 seconds" .. found it | |
| 17:54:49 | mnaser | it's actually listed but the bash for some reason doesnt find it | |
| 17:55:51 | efried | leakypipes: You around today? | |
| 17:55:58 | efried | Hoho, it's Friday | |
| 17:56:14 | andreaf | mriedem hey I'm setting up a zuul-v3 multinode job, and everything works fine apart from nova that gives me "Host is not mapped to any cell" http://logs.openstack.org/24/545724/9/check/tempest-multinode-full/1bbec81/ara/result/525a60bd-ac22-4fc4-9db7-e61fce8ac1f5/ | |
| 17:56:45 | andreaf | mriedem: I compared configs in localrc and nova and I don't see anything obvious - do you have any idea about what this could be? | |
| 17:56:52 | leakypipes | fried_rice_jfk: ues | |
| 17:56:54 | leakypipes | yes | |
| 17:57:06 | fried_rice_jfk | leakypipes: It looks like alex_xu may be right. I wrote a gabbit for it. | |
| 17:57:49 | mnaser | andreaf: im actually looking into this right ow | |
| 17:57:55 | mnaser | im seeing this issue in stable/pike | |
| 17:58:38 | andreaf | mnaser oh ok at least it's not just me :P | |
| 17:59:00 | mnaser | andreaf: i'm seeing the devstack start waiting for compute to go up at "2018-02-23 04:24:33.575" (with a 60s timeout) and the compute record get created at "2018-02-23 04:25:44.182" | |
| 17:59:31 | mnaser | with a 60 second timeout, it means that devstack gives up at 04:25:33, but the compute record actually gets created 11 seconds later | |
| 18:01:29 | fried_rice_jfk | leakypipes: Trying to figure out why. Is the order in which I construct my query supposed to not matter? | |
| 18:03:26 | fried_rice_jfk | leakypipes: Maybe query.select_from() overwrites any previous .select_from()? | |
| 18:04:16 | mnaser | so it takes 42 seconds to go from "Connecting to libvirt: qemu:///system _get_new_connection /opt/stack/old/nova/nova/virt/libvirt/host.py:366" => "Registering for lifecycle events" | |
| 18:04:38 | mnaser | and those 42 seconds are enough for devstack to timeout waiting for compute | |
| 18:05:42 | mnaser | so somehow "wrapped_conn = self._connect(self._uri, self._read_only)" takes 42 seconds here https://github.com/openstack/nova/blob/stable/pike/nova/virt/libvirt/host.py#L368 | |
| 18:07:52 | mnaser | libvirtd starts at "2018-02-23 04:23:59.741" from devstack | |
| 18:07:57 | fried_rice_jfk | leakypipes: Gah, my apologies; I was misreading my test failure. It's fine, it's cumulative as we expected. | |
| 18:08:42 | openstackgerrit | Eric Fried proposed openstack/nova master: rp: GET /resource_providers?required= |
|
| 18:08:52 | leakypipes | fried_rice_jfk: k | |
| 18:08:54 | fried_rice_jfk | leakypipes, alex_xu: Added a gabbit to prove ANDness with resources ^ | |
| 18:18:25 | mnaser | who's the person to bug for libvirt related questions | |
| 18:19:01 | mnaser | i'm seeing 1116 "device-list-properties" on service start | |
| 18:22:05 | figleaf | fried_rice_jfk: looks good! | |