| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-11 | |||
| 12:06:24 | mriedem | and we NoValidHost | |
| 12:06:26 | mriedem | like a champ | |
| 12:06:43 | dtantsur | right, maybe a wait loop until we get something in placement (hence my question about its CLI) solves it | |
| 12:06:48 | sdague | mriedem: yes, not just all of docs.o.o, also ask.o.o | |
| 12:07:01 | sdague | mriedem: that's definitely an issue | |
| 12:07:15 | sdague | mriedem: going to have to be brought up at PTG I think | |
| 12:07:27 | cdent | dtantsur: placement is so easy to curl that no one has bothered yet | |
| 12:07:45 | mriedem | sdague: complaining in -doc | |
| 12:08:01 | mriedem | cdent: there have been unmerged patches | |
| 12:08:08 | mriedem | cdent: https://review.openstack.org/#/q/project:openstack/osc-placement | |
| 12:08:10 | cdent | mriedem: yes, I know | |
| 12:08:14 | mriedem | ok | |
| 12:08:35 | cdent | i’m one of the few reviewers on those patches, and mentioned them for several months on the rp update weekly messages and finally stopped when no one was reviewing | |
| 12:08:43 | mriedem | :( | |
| 12:08:43 | cdent | because I assumed nobody cared | |
| 12:08:53 | cdent | :( is right | |
| 12:09:01 | mriedem | once we want to start integrating them into CI, people will care | |
| 12:09:03 | mriedem | well, dev people | |
| 12:10:13 | cdent | dtantsur, mriedem: so do you think you’ve gotten past at least a first hurdle with the conceptual wait loop? If so, once that’s cleared out, I’d like eventually to come back to this issue of reporting or not report inventory for nodes that have instances on them | |
| 12:10:22 | dtantsur | folks, I've dumped my/our findings on https://etherpad.openstack.org/p/nova-ironic-resource-class-migration | |
| 12:10:30 | dtantsur | I cannot keep so much in my head :) | |
| 12:10:35 | cdent | good idea | |
| 12:11:58 | cdent | if we merge this those silly html error responses (in the pastes there) will go away: https://review.openstack.org/#/c/489772/ | |
| 12:13:38 | mriedem | so like this http://logs.openstack.org/72/489772/2/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial-nv/655a0b3/logs/screen-n-cpu.txt.gz?level=TRACE#_Aug_02_12_12_16_584946 | |
| 12:14:56 | cdent | mriedem: yeah, no line feeds... | |
| 12:16:29 | mriedem | like this http://logs.openstack.org/85/490085/7/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial-nv/5835515/logs/screen-n-cpu.txt.gz?level=TRACE#_Aug_03_15_47_53_683694 | |
| 12:16:42 | cdent | right | |
| 12:23:39 | cdent | dtantsur: so currently plan is is wait and see how https://review.openstack.org/#/c/476968/ turns out? | |
| 12:24:27 | mriedem | sdague: http://lists.openstack.org/pipermail/openstack-dev/2017-August/121042.html | |
| 12:24:42 | dtantsur | cdent: this is the plan for #1. for #2 and #3, I can try fixing get_inventory indeed. | |
| 12:25:14 | cdent | dtantsur: let me know if there’s something I can help with | |
| 12:25:31 | dtantsur | sure, thanks! | |
| 12:29:26 | openstackgerrit | OpenStack Release Bot proposed openstack/nova master: Update reno for stable/pike https://review.openstack.org/492982 | |
| 12:32:55 | openstackgerrit | Dmitry Tantsur proposed openstack/nova master: DNM PoC for fixing ironic with resource classes https://review.openstack.org/492964 | |
| 12:32:57 | dtantsur | cdent: something like ^^^? | |
| 12:33:55 | dtantsur | mriedem: ^^ | |
| 12:34:52 | boolman | mriedem: ok I got it to work now, thanks | |
| 12:37:42 | dtantsur | vdrok: mind testing again with my patch above? | |
| 12:38:02 | cdent | dtantsur: yes. Was there some additional thing to do to make sure that get_inventory gets called often enough? I’m guessing (giving the log messages) that that’s not a problem? | |
| 12:38:10 | vdrok | dtantsur: ok, will do | |
| 12:38:33 | dtantsur | cdent: I'm not sure, let's see how it looks for vdrok | |
| 12:38:38 | vdrok | yeah, after instance deletion we'll have some time window when the resources reported by placement will be incorrect | |
| 12:38:47 | cdent | ✔ | |
| 12:40:55 | sdague | mriedem: I'll see if I can hack around it | |
| 12:41:24 | dtantsur | once we get custom resource classes to work with ironic, we can ask operators to upgrade. then they won't see issues with VCPU reporting.. | |
| 12:44:04 | cdent | biab | |
| 12:46:41 | bauzas | mriedem: dtantsur: could you please tl;dr the issues with ironic ? | |
| 12:46:50 | bauzas | and how I could help ? | |
| 12:47:11 | bauzas | tons of channel logs :) | |
| 12:48:26 | dtantsur | bauzas: this is the tl;dr https://etherpad.openstack.org/p/nova-ironic-resource-class-migration | |
| 12:48:47 | bauzas | excellent, thanks | |
| 12:54:06 | vdrok | dtantsur: see comment | |
| 12:54:15 | vdrok | right now requests to placement fail | |
| 12:54:21 | vdrok | because of max_unit=0 | |
| 12:54:33 | dtantsur | ugh, right | |
| 12:55:19 | openstackgerrit | Dmitry Tantsur proposed openstack/nova master: PoC for fixing ironic with resource classes https://review.openstack.org/492964 | |
| 12:55:21 | dtantsur | vdrok: please try ^^^ | |
| 12:55:25 | vdrok | yup | |
| 13:10:13 | vdrok | dtantsur: mriedem with https://review.openstack.org/492964 seems to work fine http://paste.openstack.org/show/618171/ | |
| 13:10:33 | vdrok | will now look at what's in the nova_api db | |
| 13:13:31 | figleaf | wow, it usually takes me 5 minutes to read the overnight scrollback. Today it was more like 20 | |
| 13:14:10 | figleaf | so... anything I can pitch in and help with right now? | |
| 13:16:07 | mriedem | vdrok: yeah i suppose that's why we get here now | |
| 13:16:08 | mriedem | Aug 11 13:07:03 ubuntu nova-compute[11924]: DEBUG nova.scheduler.client.report [None req-b88f01be-3920-4bc0-8708-b96b4f9e8aa7 None None] Updated inventory for 935678ef-67b2-440d-8190-875fb6dea1c6 at generation 3 {{(pid=11924) _update_inventory_attempt /opt/stack/nova/nova/scheduler/client/report.py:652}} | |
| 13:16:14 | mriedem | figleaf: https://etherpad.openstack.org/p/nova-ironic-resource-class-migration | |
| 13:16:17 | vdrok | dtantsur: mriedem what's in nova_api seems to be correct, instance_extra.flavor too, but see this http://paste.openstack.org/show/618173/ | |
| 13:16:24 | figleaf | mriedem: yeah, got that open already | |
| 13:16:35 | vdrok | namely, negative free values in compute_nodes | |
| 13:16:53 | mriedem | vdrok: that might be a latent problem? | |
| 13:17:01 | mriedem | i never look at hypervisor-stats, especially for ironic | |
| 13:17:03 | vdrok | mriedem: might be yeah | |
| 13:17:17 | vdrok | I'll try to boot another instance | |
| 13:17:26 | mriedem | vdrok: https://bugs.launchpad.net/nova/+bug/1699947 ? | |
| 13:17:27 | openstack | Launchpad bug 1699947 in OpenStack Compute (nova) "nova hypervisor-stats/hypervisor-show shows wrong resource usage for baremetal node" [Low,Confirmed] | |
| 13:17:44 | vdrok | yup, that is an old one :) | |
| 13:18:57 | vdrok | ok, scheduling seems to work fine too | |
| 13:20:51 | dtantsur | sweet! thanks for testing vdrok :) | |
| 13:21:10 | dtantsur | so, what are the next steps? wait for the CI, check that the warning is gone, then write some unit tests and merge? | |
| 13:23:13 | dtantsur | mriedem: do you think it will still try to delete the allocation? | |
| 13:25:23 | openstackgerrit | Matt Riedemann proposed openstack/nova master: doc: add superconductor up-call caveat for cross_az_attach=False https://review.openstack.org/493007 | |
| 13:25:23 | openstackgerrit | Matt Riedemann proposed openstack/nova master: doc: add another up-call caveat for cells v2 for xenapi aggregates https://review.openstack.org/493006 | |
| 13:25:25 | mriedem | dansmith: melwitt: ^ superconductor up-call limitations for the docs - i found another one today | |
| 13:25:38 | mriedem | dtantsur: you mean the inventory? | |
| 13:25:53 | dtantsur | mriedem: yes. sorry, tired already :) | |
| 13:26:02 | dtantsur | Friday is not the best day to debug Nova :) | |
| 13:26:12 | mriedem | dtantsur: no because inv_data will not be empty https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L779 | |
| 13:26:44 | dtantsur | mriedem: right, so the warning should be gone, no? we won't try to delete it, just update? | |
| 13:27:13 | mriedem | dtantsur: i'd expect to get a 409 response from placement here https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L566 | |
| 13:27:38 | mriedem | dtantsur: plus, that doesn't really fix the bug that exists in ocata already, because your fix would only bypass the delete_inventory path iff there is a node.resource_class set | |
| 13:27:54 | mriedem | dtantsur: in other words, i think there are two bugs | |
| 13:28:01 | dtantsur | the 4th problems \o/ | |
| 13:28:13 | dtantsur | should I mark it as Related-Bug then? | |
| 13:28:16 | mriedem | the one i reported about the warnings is latent, and exists in ocata | |
| 13:28:21 | mriedem | dtantsur: that would be ok probably | |
| 13:28:40 | mriedem | i.e. in ocata, if i've got a baremetal env, i'm going to see these warnings in n-cpu every 60 seconds for all nodes | |
| 13:28:51 | mriedem | for all *consumed* nodes | |
| 13:28:54 | mriedem | which is annoying | |
| 13:29:37 | dtantsur | mriedem: wait, why? we no longer return an empty inventory for valid nodes. we return an inventory with s/vcpus/vcpus_used/. I think we can backport it to Ocata even | |
| 13:30:14 | dtantsur | jaypipes: hi, you may want to join the party :) | |
| 13:30:20 | mriedem | oh i see | |