Earlier  
Posted Nick Remark
#openstack-nova - 2017-07-25
18:54:16 mriedem you can't backport counting quotas to ocata
18:54:31 mriedem but the fix might already be available, which i pointd out in the bug report and marked it incomplete since they didn't provide the version
18:54:42 sdague mriedem: sure, but fixed in master does count as fixed
18:54:55 sdague then it's a backport question
18:55:08 mriedem it's incomplete either way at this point
18:55:09 sdague but that doesn't make it not fixed
18:55:16 sdague incomplete doesn't mean that
18:55:52 sdague sure, as long as we're talking about incomplete just meaning "question back to reporter that needs an answer"
18:56:42 sdague but if it's fixed in master, then I'd say it's probably correct to comment as such and put Fix Released on it
19:07:14 openstackgerrit Matt Riedemann proposed openstack/nova master: Add oslo_concurrency=INFO to default log levels for nova-manage https://review.openstack.org/487179
19:09:32 openstackgerrit Jan Gutter proposed openstack/nova master: Add VIFHostDevice support to libvirt driver https://review.openstack.org/486426
19:18:11 mriedem sdague: dansmith: i've dumped my debug notes in https://review.openstack.org/#/c/477556/
19:18:31 mriedem we discover and map the compute node on the primary host as part of the devstack stack.sh run on the primary host,
19:18:42 mriedem we discover and map the subnodes (2 of them) after the subnodes are stacked
19:18:45 openstackgerrit Jan Gutter proposed openstack/nova master: Netronome SmartNIC Enablement https://review.openstack.org/483459
19:18:57 mriedem but it looks like when discover_hosts runs, we only discover and map the subnode-2, but miss subnode-3
19:19:10 mriedem so this might just be a latent issue in 3-node jobs
19:19:21 mriedem but makes me wonder why we don't hit this more often in 2-node jobs
19:24:53 openstackgerrit Dan Smith proposed openstack/nova master: [WIP] Add some more cellsv2 doc goodness https://review.openstack.org/487183
19:25:42 dansmith mriedem: okay I really hadn't done any 3 node thinking yet
19:25:59 mriedem dansmith: as far as i can tell, the 3rd node is being setup the same as the 2nd node
19:26:00 dansmith mriedem: are there actual production 3-node jobs that we'll break with this?
19:26:15 mriedem i don't know if there are any 3 node voting jobs, but i can dig
19:26:31 dansmith okay
19:26:34 mriedem i think we're basically getting lucky in the 2 node jobs
19:26:42 openstackgerrit Robert Ellis proposed openstack/nova master: Clarifying node_uuid usage in ironic driver. https://review.openstack.org/485803
19:27:11 mriedem for example, this is a normal 2 node job
19:27:17 mriedem we discover subnode host here
19:27:18 mriedem http://logs.openstack.org/66/483566/10/check/gate-tempest-dsvm-neutron-multinode-full-ubuntu-xenial-nv/770b47e/console.html#_2017-07-24_16_27_34_062884
19:27:21 dansmith meaning we're getting the first subnode from the main node?
19:27:23 mriedem 2017-07-24 16:27:34.062884 | + /opt/stack/new/devstack-gate/devstack-vm-gate.sh:main:L777: discover_hosts
19:28:00 mriedem http://logs.openstack.org/66/483566/10/check/gate-tempest-dsvm-neutron-multinode-full-ubuntu-xenial-nv/770b47e/logs/subnode-2/screen-n-cpu.txt.gz#_Jul_24_16_27_35_403617
19:28:00 mriedem and that subnode compute node was actually created after that
19:28:04 mriedem Jul 24 16:27:35.403617 ubuntu-xenial-2-node-osic-cloud1-disk-10046822-741313 nova-compute[1379]: INFO nova.compute.resource_tracker [None req-29fd1bd5-8730-42d5-8075-04acacfe704a None None] Compute node record created for ubuntu-xenial-2-node-osic-cloud1-disk-10046822-741313:ubuntu-xenial-2-node-osic-cloud1-disk-10046822-741313 with uuid: d09dec50-566e-41ea-adde-7ff566b63867
19:28:17 mriedem dansmith: yes the first compute node comes from the primary
19:28:44 mriedem the compute node on the primary host gets discovered as part of the primary host setup https://github.com/openstack-dev/devstack/blob/master/stack.sh#L1448
19:28:48 dansmith okay
19:29:30 mriedem so our docs say
19:29:31 mriedem "Configure and start your compute hosts. Before step 7, make sure you have compute hosts in the database by running nova service-list --binary nova-compute."
19:29:42 mriedem step 7 is running discover_hosts
19:29:55 mriedem so,
19:30:16 mriedem what we should really probably be doing is passing a variable down from devstack-gate to the tools/discover_hosts.sh script in devstack telling it how many hosts we expect to show up
19:30:19 mriedem before doing discovery
19:30:33 dansmith I thought we were specifically not supposed to do that?
19:30:39 dansmith like, I thought we had an argument about that
19:31:13 mriedem well, if you've got a slow subnode then i'm not sure what the other options are
19:31:31 mriedem we have the periodic task, but that's still a race window
19:31:34 dansmith agreed, I just thought we were told not to
19:32:27 mriedem dansmith: isn't it fun we're having the same conversation we had almost exactly 6 months ago?!
19:32:34 mriedem except i was in cabo san lucas at that time, which was more fun
19:33:52 mriedem so, i dont think this is a problem in your fleetify change which is the key point
19:34:00 dansmith mriedem: do these graphs do anything for you? http://docs-draft.openstack.org/83/487183/1/check/gate-nova-docs-ubuntu-xenial/ef53873//doc/build/html/user/cellsv2_layout.html
19:34:00 mriedem it's just a latent thing we aren't handling well in our infra
19:34:01 dansmith mostly the second one
19:34:11 dansmith okay
19:34:44 openstackgerrit Ildiko Vancsa proposed openstack/nova master: Translate the return value of attachment_create and _update https://review.openstack.org/486194
19:35:10 mriedem dansmith: looks pretty good
19:35:38 mriedem the first graph is kind of a webby mess
19:35:40 dansmith wish I could fix a few visual aberrations on it, but it took a lot of screwing around to make it look this good
19:36:11 dansmith I spent less time on the first one I can muck some more
19:40:06 sdague mriedem: ok, so have you confirmed if the 3 node is voting yet?
19:40:10 openstackgerrit Moshe Levi proposed openstack/nova master: hardware offload support for openvswitch https://review.openstack.org/398265
19:40:42 sdague mriedem: I thought we were doing discover hosts at the end of the d-g run?
19:41:55 mriedem sdague: we are,
19:42:07 mriedem but the compute node getting created is asynchronous to that
19:42:18 mriedem the 3-node job could be slowing down the controller services just enough to hit the latent window
19:42:36 sdague because the startup of nova-compute takes that long?
19:43:14 sdague so the race is that nova-compute service start doesn't make it to the db before discover hosts runs?
19:44:59 mriedem yes
19:45:19 mriedem +1 on https://review.openstack.org/#/c/477556/ and my debug notes are all in there
19:46:36 sdague so... related, what's the deal with the stack trace here - http://logs.openstack.org/56/477556/5/experimental/gate-tempest-dsvm-neutron-dvr-ha-multinode-full-ubuntu-xenial-nv/432c235/logs/subnode-3/screen-n-cpu.txt.gz#_Jul_25_15_06_55_309283
19:47:23 mriedem that's the thing where the libvirt starts up and tries to enable itself
19:47:54 sdague ok, so we're going to stacktrace on every clean start
19:48:06 mriedem https://github.com/openstack/nova/blob/master/nova/virt/libvirt/driver.py#L3563
19:48:09 mriedem sdague: that's been around
19:48:11 mriedem it's not a result of this change
19:48:19 sdague mriedem: sure
19:48:23 sdague it's just not good
19:48:37 mriedem yeah, i don't like it either
19:48:42 sdague ok, http://logs.openstack.org/56/477556/5/experimental/gate-tempest-dsvm-neutron-dvr-ha-multinode-full-ubuntu-xenial-nv/432c235/logs/subnode-3/screen-n-cpu.txt.gz#_Jul_25_15_07_02_323379 is where the compute node is built, that's about 7 seconds later
19:48:54 mriedem yes
19:49:51 mriedem i'm not sure why these stacktrace
19:50:21 mriedem i guess because of https://github.com/openstack/nova/blob/master/nova/virt/libvirt/driver.py#L3567 ?
19:50:33 mriedem so it hits the generic Exception block as ComputeHostNotFound_Remote?
19:50:39 openstackgerrit Dan Smith proposed openstack/nova master: [WIP] Add some more cellsv2 doc goodness https://review.openstack.org/487183
19:52:56 sdague mriedem: yeh
19:59:42 sdague mriedem: do we have a way of telling that nova-compute is ready. Is that resource tracker at 07_02 the event we need?
20:00:26 openstackgerrit Merged openstack/nova master: API ref: associate floating IP requires Active status https://review.openstack.org/363642
20:00:48 sdague I'm trying to think about who should wait for what to get us there. Is there something we could tell from the subnode easily about it being ready so we knew we were in the clear when stack.sh finished?
20:01:18 mriedem sdague: our docs tell you to run 'nova service-list --binary nova-compute' and make sure the compute shows up before you run discover_hosts
20:02:37 mriedem because that API iterates all cells and gathers up the services running in them
20:02:42 mriedem so the host mapping isn't required for that
20:03:24 sdague so we could put a flag in to wait for compute
20:03:29 sdague and run: nova service-list --host `hostname` --binary nova-compute
20:03:43 sdague on the child until it is true?
20:08:00 dansmith not on the child, on the main node
20:08:41 mriedem right we run discover_hosts from the primary
20:08:48 mriedem b/c it needs to get to the api db to find the cell mappings
20:08:58 mriedem and the compute nodes don't have access to that
20:09:00 dansmith well and you need the main node to do the waiting

Earlier   Later