Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-11
12:05:01 aarefiev dtantsur: but we report empty inventory, right https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L741
12:05:02 mriedem regardless of the resource class stuff
12:05:17 dtantsur mriedem: we used to wait for 'nova hypervisor-stats' to show available CPUS > 0
12:05:23 mriedem dtantsur: and you're not now?
12:05:48 dtantsur mriedem: that patch tries not reporting VCPU any more
12:06:00 dtantsur maybe it's a bit premature
12:06:09 mriedem ok, so we don't wait for the compute node to be registered in nova-compute,
12:06:14 mriedem and try to create the instance,
12:06:19 mriedem and since we didn't wait, we don't have the resource class created yet
12:06:24 mriedem and we NoValidHost
12:06:26 mriedem like a champ
12:06:43 dtantsur right, maybe a wait loop until we get something in placement (hence my question about its CLI) solves it
12:06:48 sdague mriedem: yes, not just all of docs.o.o, also ask.o.o
12:07:01 sdague mriedem: that's definitely an issue
12:07:15 sdague mriedem: going to have to be brought up at PTG I think
12:07:27 cdent dtantsur: placement is so easy to curl that no one has bothered yet
12:07:45 mriedem sdague: complaining in -doc
12:08:01 mriedem cdent: there have been unmerged patches
12:08:08 mriedem cdent: https://review.openstack.org/#/q/project:openstack/osc-placement
12:08:10 cdent mriedem: yes, I know
12:08:14 mriedem ok
12:08:35 cdent i’m one of the few reviewers on those patches, and mentioned them for several months on the rp update weekly messages and finally stopped when no one was reviewing
12:08:43 cdent because I assumed nobody cared
12:08:43 mriedem :(
12:08:53 cdent :( is right
12:09:01 mriedem once we want to start integrating them into CI, people will care
12:09:03 mriedem well, dev people
12:10:13 cdent dtantsur, mriedem: so do you think you’ve gotten past at least a first hurdle with the conceptual wait loop? If so, once that’s cleared out, I’d like eventually to come back to this issue of reporting or not report inventory for nodes that have instances on them
12:10:22 dtantsur folks, I've dumped my/our findings on https://etherpad.openstack.org/p/nova-ironic-resource-class-migration
12:10:30 dtantsur I cannot keep so much in my head :)
12:10:35 cdent good idea
12:11:58 cdent if we merge this those silly html error responses (in the pastes there) will go away: https://review.openstack.org/#/c/489772/
12:13:38 mriedem so like this http://logs.openstack.org/72/489772/2/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial-nv/655a0b3/logs/screen-n-cpu.txt.gz?level=TRACE#_Aug_02_12_12_16_584946
12:14:56 cdent mriedem: yeah, no line feeds...
12:16:29 mriedem like this http://logs.openstack.org/85/490085/7/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial-nv/5835515/logs/screen-n-cpu.txt.gz?level=TRACE#_Aug_03_15_47_53_683694
12:16:42 cdent right
12:23:39 cdent dtantsur: so currently plan is is wait and see how https://review.openstack.org/#/c/476968/ turns out?
12:24:27 mriedem sdague: http://lists.openstack.org/pipermail/openstack-dev/2017-August/121042.html
12:24:42 dtantsur cdent: this is the plan for #1. for #2 and #3, I can try fixing get_inventory indeed.
12:25:14 cdent dtantsur: let me know if there’s something I can help with
12:25:31 dtantsur sure, thanks!
12:29:26 openstackgerrit OpenStack Release Bot proposed openstack/nova master: Update reno for stable/pike https://review.openstack.org/492982
12:32:55 openstackgerrit Dmitry Tantsur proposed openstack/nova master: DNM PoC for fixing ironic with resource classes https://review.openstack.org/492964
12:32:57 dtantsur cdent: something like ^^^?
12:33:55 dtantsur mriedem: ^^
12:34:52 boolman mriedem: ok I got it to work now, thanks
12:37:42 dtantsur vdrok: mind testing again with my patch above?
12:38:02 cdent dtantsur: yes. Was there some additional thing to do to make sure that get_inventory gets called often enough? I’m guessing (giving the log messages) that that’s not a problem?
12:38:10 vdrok dtantsur: ok, will do
12:38:33 dtantsur cdent: I'm not sure, let's see how it looks for vdrok
12:38:38 vdrok yeah, after instance deletion we'll have some time window when the resources reported by placement will be incorrect
12:38:47 cdent ✔
12:40:55 sdague mriedem: I'll see if I can hack around it
12:41:24 dtantsur once we get custom resource classes to work with ironic, we can ask operators to upgrade. then they won't see issues with VCPU reporting..
12:44:04 cdent biab
12:46:41 bauzas mriedem: dtantsur: could you please tl;dr the issues with ironic ?
12:46:50 bauzas and how I could help ?
12:47:11 bauzas tons of channel logs :)
12:48:26 dtantsur bauzas: this is the tl;dr https://etherpad.openstack.org/p/nova-ironic-resource-class-migration
12:48:47 bauzas excellent, thanks
12:54:06 vdrok dtantsur: see comment
12:54:15 vdrok right now requests to placement fail
12:54:21 vdrok because of max_unit=0
12:54:33 dtantsur ugh, right
12:55:19 openstackgerrit Dmitry Tantsur proposed openstack/nova master: PoC for fixing ironic with resource classes https://review.openstack.org/492964
12:55:21 dtantsur vdrok: please try ^^^
12:55:25 vdrok yup
13:10:13 vdrok dtantsur: mriedem with https://review.openstack.org/492964 seems to work fine http://paste.openstack.org/show/618171/
13:10:33 vdrok will now look at what's in the nova_api db
13:13:31 figleaf wow, it usually takes me 5 minutes to read the overnight scrollback. Today it was more like 20
13:14:10 figleaf so... anything I can pitch in and help with right now?
13:16:07 mriedem vdrok: yeah i suppose that's why we get here now
13:16:08 mriedem Aug 11 13:07:03 ubuntu nova-compute[11924]: DEBUG nova.scheduler.client.report [None req-b88f01be-3920-4bc0-8708-b96b4f9e8aa7 None None] Updated inventory for 935678ef-67b2-440d-8190-875fb6dea1c6 at generation 3 {{(pid=11924) _update_inventory_attempt /opt/stack/nova/nova/scheduler/client/report.py:652}}
13:16:14 mriedem figleaf: https://etherpad.openstack.org/p/nova-ironic-resource-class-migration
13:16:17 vdrok dtantsur: mriedem what's in nova_api seems to be correct, instance_extra.flavor too, but see this http://paste.openstack.org/show/618173/
13:16:24 figleaf mriedem: yeah, got that open already
13:16:35 vdrok namely, negative free values in compute_nodes
13:16:53 mriedem vdrok: that might be a latent problem?
13:17:01 mriedem i never look at hypervisor-stats, especially for ironic
13:17:03 vdrok mriedem: might be yeah
13:17:17 vdrok I'll try to boot another instance
13:17:26 mriedem vdrok: https://bugs.launchpad.net/nova/+bug/1699947 ?
13:17:27 openstack Launchpad bug 1699947 in OpenStack Compute (nova) "nova hypervisor-stats/hypervisor-show shows wrong resource usage for baremetal node" [Low,Confirmed]
13:17:44 vdrok yup, that is an old one :)
13:18:57 vdrok ok, scheduling seems to work fine too
13:20:51 dtantsur sweet! thanks for testing vdrok :)
13:21:10 dtantsur so, what are the next steps? wait for the CI, check that the warning is gone, then write some unit tests and merge?
13:23:13 dtantsur mriedem: do you think it will still try to delete the allocation?
13:25:23 openstackgerrit Matt Riedemann proposed openstack/nova master: doc: add another up-call caveat for cells v2 for xenapi aggregates https://review.openstack.org/493006
13:25:23 openstackgerrit Matt Riedemann proposed openstack/nova master: doc: add superconductor up-call caveat for cross_az_attach=False https://review.openstack.org/493007
13:25:25 mriedem dansmith: melwitt: ^ superconductor up-call limitations for the docs - i found another one today
13:25:38 mriedem dtantsur: you mean the inventory?
13:25:53 dtantsur mriedem: yes. sorry, tired already :)
13:26:02 dtantsur Friday is not the best day to debug Nova :)
13:26:12 mriedem dtantsur: no because inv_data will not be empty https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L779
13:26:44 dtantsur mriedem: right, so the warning should be gone, no? we won't try to delete it, just update?
13:27:13 mriedem dtantsur: i'd expect to get a 409 response from placement here https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L566
13:27:38 mriedem dtantsur: plus, that doesn't really fix the bug that exists in ocata already, because your fix would only bypass the delete_inventory path iff there is a node.resource_class set
13:27:54 mriedem dtantsur: in other words, i think there are two bugs
13:28:01 dtantsur the 4th problems \o/

Earlier   Later