| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-07-25 | |||
| 18:36:19 | mriedem | ildikov: sorry i'm in a meeting and doing a few things at once and i don't have all of your changes in my head right now, | |
| 18:36:28 | mriedem | but in general, we shouldn't have to touch nova/tests/functional for the new stuff, | |
| 18:36:49 | mriedem | since that is mostly all for api samples and notification samples, which are stubbing out cinder in specific ways for the old flows | |
| 18:37:05 | mriedem | i mostly care about test coverage for the *new* flows using unit tests | |
| 18:37:19 | mriedem | we can worry about changing the various fixtures and functional tests over when we actually drop the *old* flows | |
| 18:37:23 | mriedem | which is not going to be anytime soon | |
| 18:37:39 | ildikov | mriedem: I figured out the swap stuff already and the rest is simple, the test failure is a mock and inheritance issue, and has nothing to do with Cinder | |
| 18:38:39 | ildikov | mriedem: but anyway, if you don't want changes there I will remove it and I will add mocks to the places which fails due to having the highest service version but no new Cinder calls available to switch back to the old flow and these tests in a follow up patch | |
| 18:39:30 | ildikov | mriedem: anyway, sorry for eating up this much of your time, I will go and do the updates we agreed on and then we can check if there's anything else to fix | |
| 18:39:49 | mriedem | ildikov: just leave the tests you have then in the api change, i can look into them at some point | |
| 18:40:13 | ildikov | mriedem: ok | |
| 18:40:19 | ildikov | mriedem: thanks | |
| 18:50:45 | sdague | melwitt: if you could comment on the bug, and move it to Fix Released, that would be cool | |
| 18:53:07 | sdague | mriedem: back for a bit, you look through those results yet? | |
| 18:53:26 | sdague | also, anyone, this is a pretty easy deprecation of a conf variable - https://review.openstack.org/#/c/486623/ | |
| 18:54:10 | mriedem | sdague: melwitt: the quotas bug for ocata pointed out earlier is not necessarily fix released | |
| 18:54:16 | mriedem | you can't backport counting quotas to ocata | |
| 18:54:31 | mriedem | but the fix might already be available, which i pointd out in the bug report and marked it incomplete since they didn't provide the version | |
| 18:54:42 | sdague | mriedem: sure, but fixed in master does count as fixed | |
| 18:54:55 | sdague | then it's a backport question | |
| 18:55:08 | mriedem | it's incomplete either way at this point | |
| 18:55:09 | sdague | but that doesn't make it not fixed | |
| 18:55:16 | sdague | incomplete doesn't mean that | |
| 18:55:52 | sdague | sure, as long as we're talking about incomplete just meaning "question back to reporter that needs an answer" | |
| 18:56:42 | sdague | but if it's fixed in master, then I'd say it's probably correct to comment as such and put Fix Released on it | |
| 19:07:14 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Add oslo_concurrency=INFO to default log levels for nova-manage https://review.openstack.org/487179 | |
| 19:09:32 | openstackgerrit | Jan Gutter proposed openstack/nova master: Add VIFHostDevice support to libvirt driver https://review.openstack.org/486426 | |
| 19:18:11 | mriedem | sdague: dansmith: i've dumped my debug notes in https://review.openstack.org/#/c/477556/ | |
| 19:18:31 | mriedem | we discover and map the compute node on the primary host as part of the devstack stack.sh run on the primary host, | |
| 19:18:42 | mriedem | we discover and map the subnodes (2 of them) after the subnodes are stacked | |
| 19:18:45 | openstackgerrit | Jan Gutter proposed openstack/nova master: Netronome SmartNIC Enablement https://review.openstack.org/483459 | |
| 19:18:57 | mriedem | but it looks like when discover_hosts runs, we only discover and map the subnode-2, but miss subnode-3 | |
| 19:19:10 | mriedem | so this might just be a latent issue in 3-node jobs | |
| 19:19:21 | mriedem | but makes me wonder why we don't hit this more often in 2-node jobs | |
| 19:24:53 | openstackgerrit | Dan Smith proposed openstack/nova master: [WIP] Add some more cellsv2 doc goodness https://review.openstack.org/487183 | |
| 19:25:42 | dansmith | mriedem: okay I really hadn't done any 3 node thinking yet | |
| 19:25:59 | mriedem | dansmith: as far as i can tell, the 3rd node is being setup the same as the 2nd node | |
| 19:26:00 | dansmith | mriedem: are there actual production 3-node jobs that we'll break with this? | |
| 19:26:15 | mriedem | i don't know if there are any 3 node voting jobs, but i can dig | |
| 19:26:31 | dansmith | okay | |
| 19:26:34 | mriedem | i think we're basically getting lucky in the 2 node jobs | |
| 19:26:42 | openstackgerrit | Robert Ellis proposed openstack/nova master: Clarifying node_uuid usage in ironic driver. https://review.openstack.org/485803 | |
| 19:27:11 | mriedem | for example, this is a normal 2 node job | |
| 19:27:17 | mriedem | we discover subnode host here | |
| 19:27:18 | mriedem | http://logs.openstack.org/66/483566/10/check/gate-tempest-dsvm-neutron-multinode-full-ubuntu-xenial-nv/770b47e/console.html#_2017-07-24_16_27_34_062884 | |
| 19:27:21 | dansmith | meaning we're getting the first subnode from the main node? | |
| 19:27:23 | mriedem | 2017-07-24 16:27:34.062884 | + /opt/stack/new/devstack-gate/devstack-vm-gate.sh:main:L777: discover_hosts | |
| 19:28:00 | mriedem | http://logs.openstack.org/66/483566/10/check/gate-tempest-dsvm-neutron-multinode-full-ubuntu-xenial-nv/770b47e/logs/subnode-2/screen-n-cpu.txt.gz#_Jul_24_16_27_35_403617 | |
| 19:28:00 | mriedem | and that subnode compute node was actually created after that | |
| 19:28:04 | mriedem | Jul 24 16:27:35.403617 ubuntu-xenial-2-node-osic-cloud1-disk-10046822-741313 nova-compute[1379]: INFO nova.compute.resource_tracker [None req-29fd1bd5-8730-42d5-8075-04acacfe704a None None] Compute node record created for ubuntu-xenial-2-node-osic-cloud1-disk-10046822-741313:ubuntu-xenial-2-node-osic-cloud1-disk-10046822-741313 with uuid: d09dec50-566e-41ea-adde-7ff566b63867 | |
| 19:28:17 | mriedem | dansmith: yes the first compute node comes from the primary | |
| 19:28:44 | mriedem | the compute node on the primary host gets discovered as part of the primary host setup https://github.com/openstack-dev/devstack/blob/master/stack.sh#L1448 | |
| 19:28:48 | dansmith | okay | |
| 19:29:30 | mriedem | so our docs say | |
| 19:29:31 | mriedem | "Configure and start your compute hosts. Before step 7, make sure you have compute hosts in the database by running nova service-list --binary nova-compute." | |
| 19:29:42 | mriedem | step 7 is running discover_hosts | |
| 19:29:55 | mriedem | so, | |
| 19:30:16 | mriedem | what we should really probably be doing is passing a variable down from devstack-gate to the tools/discover_hosts.sh script in devstack telling it how many hosts we expect to show up | |
| 19:30:19 | mriedem | before doing discovery | |
| 19:30:33 | dansmith | I thought we were specifically not supposed to do that? | |
| 19:30:39 | dansmith | like, I thought we had an argument about that | |
| 19:31:13 | mriedem | well, if you've got a slow subnode then i'm not sure what the other options are | |
| 19:31:31 | mriedem | we have the periodic task, but that's still a race window | |
| 19:31:34 | dansmith | agreed, I just thought we were told not to | |
| 19:32:27 | mriedem | dansmith: isn't it fun we're having the same conversation we had almost exactly 6 months ago?! | |
| 19:32:34 | mriedem | except i was in cabo san lucas at that time, which was more fun | |
| 19:33:52 | mriedem | so, i dont think this is a problem in your fleetify change which is the key point | |
| 19:34:00 | dansmith | mriedem: do these graphs do anything for you? http://docs-draft.openstack.org/83/487183/1/check/gate-nova-docs-ubuntu-xenial/ef53873//doc/build/html/user/cellsv2_layout.html | |
| 19:34:00 | mriedem | it's just a latent thing we aren't handling well in our infra | |
| 19:34:01 | dansmith | mostly the second one | |
| 19:34:11 | dansmith | okay | |
| 19:34:44 | openstackgerrit | Ildiko Vancsa proposed openstack/nova master: Translate the return value of attachment_create and _update https://review.openstack.org/486194 | |
| 19:35:10 | mriedem | dansmith: looks pretty good | |
| 19:35:38 | mriedem | the first graph is kind of a webby mess | |
| 19:35:40 | dansmith | wish I could fix a few visual aberrations on it, but it took a lot of screwing around to make it look this good | |
| 19:36:11 | dansmith | I spent less time on the first one I can muck some more | |
| 19:40:06 | sdague | mriedem: ok, so have you confirmed if the 3 node is voting yet? | |
| 19:40:10 | openstackgerrit | Moshe Levi proposed openstack/nova master: hardware offload support for openvswitch https://review.openstack.org/398265 | |
| 19:40:42 | sdague | mriedem: I thought we were doing discover hosts at the end of the d-g run? | |
| 19:41:55 | mriedem | sdague: we are, | |
| 19:42:07 | mriedem | but the compute node getting created is asynchronous to that | |
| 19:42:18 | mriedem | the 3-node job could be slowing down the controller services just enough to hit the latent window | |
| 19:42:36 | sdague | because the startup of nova-compute takes that long? | |
| 19:43:14 | sdague | so the race is that nova-compute service start doesn't make it to the db before discover hosts runs? | |
| 19:44:59 | mriedem | yes | |
| 19:45:19 | mriedem | +1 on https://review.openstack.org/#/c/477556/ and my debug notes are all in there | |
| 19:46:36 | sdague | so... related, what's the deal with the stack trace here - http://logs.openstack.org/56/477556/5/experimental/gate-tempest-dsvm-neutron-dvr-ha-multinode-full-ubuntu-xenial-nv/432c235/logs/subnode-3/screen-n-cpu.txt.gz#_Jul_25_15_06_55_309283 | |
| 19:47:23 | mriedem | that's the thing where the libvirt starts up and tries to enable itself | |
| 19:47:54 | sdague | ok, so we're going to stacktrace on every clean start | |
| 19:48:06 | mriedem | https://github.com/openstack/nova/blob/master/nova/virt/libvirt/driver.py#L3563 | |
| 19:48:09 | mriedem | sdague: that's been around | |
| 19:48:11 | mriedem | it's not a result of this change | |
| 19:48:19 | sdague | mriedem: sure | |
| 19:48:23 | sdague | it's just not good | |
| 19:48:37 | mriedem | yeah, i don't like it either | |
| 19:48:42 | sdague | ok, http://logs.openstack.org/56/477556/5/experimental/gate-tempest-dsvm-neutron-dvr-ha-multinode-full-ubuntu-xenial-nv/432c235/logs/subnode-3/screen-n-cpu.txt.gz#_Jul_25_15_07_02_323379 is where the compute node is built, that's about 7 seconds later | |
| 19:48:54 | mriedem | yes | |
| 19:49:51 | mriedem | i'm not sure why these stacktrace | |
| 19:50:21 | mriedem | i guess because of https://github.com/openstack/nova/blob/master/nova/virt/libvirt/driver.py#L3567 ? | |
| 19:50:33 | mriedem | so it hits the generic Exception block as ComputeHostNotFound_Remote? | |