| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-02-23 | |||
| 16:59:17 | cfriesen | so you'd normally run with it strict, but if you need to take down a compute node you can set best-effort, move everything, then turn it back off | |
| 16:59:41 | cfriesen | otherwise if you've got hard-affinity you can't migrate any of them | |
| 17:00:04 | mriedem | which is why people want a force flag on evacuate, live migrate (and now cold migrate) | |
| 17:00:06 | leakypipes | mriedem: and the late affinity check is (the only?) remaining upcall from a cell to API, no? | |
| 17:00:15 | mriedem | leakypipes: hells no | |
| 17:00:32 | leakypipes | mriedem: it's not an upcall? | |
| 17:00:35 | mriedem | https://docs.openstack.org/nova/latest/user/cellsv2-layout.html#operations-requiring-upcalls | |
| 17:00:38 | mriedem | it is an upcall | |
| 17:00:42 | mriedem | but it's not the only one | |
| 17:00:44 | leakypipes | oh, not the only reminaing.. | |
| 17:00:46 | mriedem | we've got aggregates too | |
| 17:01:04 | mriedem | so aggregates and affinity are the remaining upcall issues | |
| 17:01:21 | mriedem | but right now there are 2 each | |
| 17:01:24 | mriedem | so 4 upcall issues | |
| 17:01:33 | leakypipes | mriedem: then you'll LOVE my aggregate affinity spec! :) now with moar AFFINITY and moar AGGREGATES! | |
| 17:01:48 | mriedem | leakypipes: i already said 'upcall upcall upcall' on that spec several times :) | |
| 17:01:54 | leakypipes | I know :) | |
| 17:02:20 | mriedem | i think i also hedged with something like, 'but we already do this in a few other places so people already have to rely on it, so maybe another log on the fire doesn't kill us' | |
| 17:02:55 | cfriesen | On a totally different topic...has anyone ever heard of nova allocating duplicate network interfaces? (So the user boots while asking for 2 network interfaces, and nova allocates two ports on each network.) | |
| 17:03:08 | mriedem | yes | |
| 17:03:10 | mriedem | that's old news | |
| 17:03:25 | mriedem | tempest has a test for it also i think | |
| 17:03:31 | mriedem | been around since juno? | |
| 17:05:28 | cfriesen | I mean nova is allocating twice as many as were asked for. | |
| 17:05:47 | mriedem | double your pleasure | |
| 17:05:49 | mriedem | idk, going to lunch | |
| 17:14:43 | efried | Greetings from JFK airport | |
| 17:27:30 | mnaser | ok | |
| 17:27:33 | mnaser | im convinced grenade is broken | |
| 17:27:43 | mnaser | for stable/pike | |
| 17:28:05 | efried | Isn't that mnaser guy known for being johnny-on-the-spot for grenade fixes? | |
| 17:28:16 | mnaser | only when i have to :( | |
| 17:28:29 | mnaser | Host 'ubuntu-xenial-rax-dfw-0002683360' is not mapped to any cell | |
| 17:28:36 | mnaser | we keep getting this in multinode | |
| 17:29:49 | efried | It would seem odd that cell discovery isn't being run. Like, nothing would ever work. | |
| 17:30:00 | efried | And that's pretty much the only thing I know about cells. | |
| 17:30:40 | mnaser | efried: indeed seems to be the case. looks like there was a change to add 'CELLSV2_SETUP=singleconductor' in there, so not sure if that might have affected it | |
| 17:31:13 | efried | mnaser: You're going to have something of a hard time finding a core today, but I'll +1 your fix :) | |
| 17:31:31 | mnaser | efried: looks like there's only 3 cores for grenade too.. | |
| 17:32:49 | efried | Looks like qa-release is included by inheritance. | |
| 17:33:02 | efried | So seven | |
| 17:35:12 | mnaser | ok looks like this runs => nova-manage cell_v2 simple_cell_setup --transport-url rabbit://stackrabbit:secretrabbit@10.209.130.218:5672/ | |
| 17:35:39 | mnaser | but discover_hosts is never called | |
| 17:36:22 | efried | And simple_cell_setup doesn't run discovery itself? | |
| 17:36:52 | efried | I remember having to fix this around pike timeframe. | |
| 17:37:02 | mnaser | efried: going through the code it looks like it does call _map_cell_and_hosts() | |
| 17:47:51 | openstack | Launchpad bug 1708039 in devstack "gate-grenade-dsvm-neutron-multinode-ubuntu-xenial fails with "No host-to-cell mapping found for selected host"" [Medium,Fix released] - Assigned to Sean Dague (sdague) | |
| 17:47:51 | mnaser | sigh https://bugs.launchpad.net/grenade/+bug/1708039 looks like it was 'supposed' to be fixed | |
| 17:49:09 | mnaser | it looks like it regressed | |
| 17:49:12 | mnaser | and its all stable/pike hits | |
| 17:54:40 | mnaser | "Didn't find service registered by hostname after 60 seconds" .. found it | |
| 17:54:49 | mnaser | it's actually listed but the bash for some reason doesnt find it | |
| 17:55:51 | efried | leakypipes: You around today? | |
| 17:55:58 | efried | Hoho, it's Friday | |
| 17:56:14 | andreaf | mriedem hey I'm setting up a zuul-v3 multinode job, and everything works fine apart from nova that gives me "Host is not mapped to any cell" http://logs.openstack.org/24/545724/9/check/tempest-multinode-full/1bbec81/ara/result/525a60bd-ac22-4fc4-9db7-e61fce8ac1f5/ | |
| 17:56:45 | andreaf | mriedem: I compared configs in localrc and nova and I don't see anything obvious - do you have any idea about what this could be? | |
| 17:56:52 | leakypipes | fried_rice_jfk: ues | |
| 17:56:54 | leakypipes | yes | |
| 17:57:06 | fried_rice_jfk | leakypipes: It looks like alex_xu may be right. I wrote a gabbit for it. | |
| 17:57:49 | mnaser | andreaf: im actually looking into this right ow | |
| 17:57:55 | mnaser | im seeing this issue in stable/pike | |
| 17:58:38 | andreaf | mnaser oh ok at least it's not just me :P | |
| 17:59:00 | mnaser | andreaf: i'm seeing the devstack start waiting for compute to go up at "2018-02-23 04:24:33.575" (with a 60s timeout) and the compute record get created at "2018-02-23 04:25:44.182" | |
| 17:59:31 | mnaser | with a 60 second timeout, it means that devstack gives up at 04:25:33, but the compute record actually gets created 11 seconds later | |
| 18:01:29 | fried_rice_jfk | leakypipes: Trying to figure out why. Is the order in which I construct my query supposed to not matter? | |
| 18:03:26 | fried_rice_jfk | leakypipes: Maybe query.select_from() overwrites any previous .select_from()? | |
| 18:04:16 | mnaser | so it takes 42 seconds to go from "Connecting to libvirt: qemu:///system _get_new_connection /opt/stack/old/nova/nova/virt/libvirt/host.py:366" => "Registering for lifecycle events" | |
| 18:04:38 | mnaser | and those 42 seconds are enough for devstack to timeout waiting for compute | |
| 18:05:42 | mnaser | so somehow "wrapped_conn = self._connect(self._uri, self._read_only)" takes 42 seconds here https://github.com/openstack/nova/blob/stable/pike/nova/virt/libvirt/host.py#L368 | |
| 18:07:52 | mnaser | libvirtd starts at "2018-02-23 04:23:59.741" from devstack | |
| 18:07:57 | fried_rice_jfk | leakypipes: Gah, my apologies; I was misreading my test failure. It's fine, it's cumulative as we expected. | |
| 18:08:42 | openstackgerrit | Eric Fried proposed openstack/nova master: rp: GET /resource_providers?required= |
|
| 18:08:52 | leakypipes | fried_rice_jfk: k | |
| 18:08:54 | fried_rice_jfk | leakypipes, alex_xu: Added a gabbit to prove ANDness with resources ^ | |
| 18:18:25 | mnaser | who's the person to bug for libvirt related questions | |
| 18:19:01 | mnaser | i'm seeing 1116 "device-list-properties" on service start | |
| 18:22:05 | figleaf | fried_rice_jfk: looks good! | |
| 18:22:22 | fried_rice_jfk | Thanks figleaf | |
| 18:25:06 | openstackgerrit | Eric Berglund proposed openstack/nova master: PowerVM Driver: vSCSI volume driver https://review.openstack.org/526094 | |
| 19:17:12 | mnaser | o/ | |
| 19:20:33 | mriedem | andreaf: mnaser: maybe on a slow test node, | |
| 19:20:58 | mriedem | devstack times out waiting for the compute host to get mapped to the cell, which won't happen until after the compute node record gets auto-created when the nova-compute service starts up | |
| 19:21:29 | mnaser | mriedem: exactly, but libvirt takes 42 seconds to start up and seems to be doing a ton of commands before starting up | |
| 19:21:38 | mnaser | sorry, not 42 | |
| 19:21:40 | mnaser | it takes 1m22s | |
| 19:22:01 | mnaser | i wonder if we should add wait_for_libvirt :\ | |
| 19:22:26 | mnaser | seeing a tooon of "Send command '{"execute":"device-list-properties","arguments":{"typename":"kvm-pci-assign"},"id":"libvirt-19"}' for write with FD -1" | |
| 19:22:50 | mriedem | not sure why libvirt would take a long time to startup in pike and not ocata...or is it the old side of grenade failing, meaning it's ocata? | |
| 19:23:04 | mnaser | mriedem: old side is failing | |
| 19:23:08 | mriedem | ok so that's ocata | |
| 19:23:19 | mriedem | but not sure why anything would have changed there wrt libvirt | |
| 19:25:39 | mnaser | looks like uca published libvirt-bin_2.5.0-3ubuntu5.6~cloud2_amd64.deb on 9th of february | |
| 19:26:06 | mnaser | prior to that libvirt-bin_2.5.0-3ubuntu5.6~cloud1_amd64.deb was released 23 jan 2018 | |
| 19:26:11 | mnaser | and before that.. nothing in 2.x series | |
| 19:27:29 | mriedem | ocata and pike are both using the ocata UCA https://github.com/openstack-dev/devstack/blob/stable/ocata/tools/fixup_stuff.sh#L84 | |
| 19:27:38 | mriedem | https://github.com/openstack-dev/devstack/blob/stable/pike/tools/fixup_stuff.sh#L87 | |
| 19:29:46 | mnaser | mriedem: interesting, the issue doesnt seem to be there with the master which uses 3.6.0 | |
| 19:31:32 | mnaser | mriedem: i dunno, i feel like bumping that timeout to 120s, but that might mask regressions with nova taking too long to register for whatever reason | |
| 19:31:53 | mnaser | i guess i can write a wait_for_libvirt | |
| 19:33:38 | mriedem | i'd be ok with i think bumping the timeout on pike/ocata | |