Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-11
12:25:31 dtantsur sure, thanks!
12:29:26 openstackgerrit OpenStack Release Bot proposed openstack/nova master: Update reno for stable/pike https://review.openstack.org/492982
12:32:55 openstackgerrit Dmitry Tantsur proposed openstack/nova master: DNM PoC for fixing ironic with resource classes https://review.openstack.org/492964
12:32:57 dtantsur cdent: something like ^^^?
12:33:55 dtantsur mriedem: ^^
12:34:52 boolman mriedem: ok I got it to work now, thanks
12:37:42 dtantsur vdrok: mind testing again with my patch above?
12:38:02 cdent dtantsur: yes. Was there some additional thing to do to make sure that get_inventory gets called often enough? I’m guessing (giving the log messages) that that’s not a problem?
12:38:10 vdrok dtantsur: ok, will do
12:38:33 dtantsur cdent: I'm not sure, let's see how it looks for vdrok
12:38:38 vdrok yeah, after instance deletion we'll have some time window when the resources reported by placement will be incorrect
12:38:47 cdent ✔
12:40:55 sdague mriedem: I'll see if I can hack around it
12:41:24 dtantsur once we get custom resource classes to work with ironic, we can ask operators to upgrade. then they won't see issues with VCPU reporting..
12:44:04 cdent biab
12:46:41 bauzas mriedem: dtantsur: could you please tl;dr the issues with ironic ?
12:46:50 bauzas and how I could help ?
12:47:11 bauzas tons of channel logs :)
12:48:26 dtantsur bauzas: this is the tl;dr https://etherpad.openstack.org/p/nova-ironic-resource-class-migration
12:48:47 bauzas excellent, thanks
12:54:06 vdrok dtantsur: see comment
12:54:15 vdrok right now requests to placement fail
12:54:21 vdrok because of max_unit=0
12:54:33 dtantsur ugh, right
12:55:19 openstackgerrit Dmitry Tantsur proposed openstack/nova master: PoC for fixing ironic with resource classes https://review.openstack.org/492964
12:55:21 dtantsur vdrok: please try ^^^
12:55:25 vdrok yup
13:10:13 vdrok dtantsur: mriedem with https://review.openstack.org/492964 seems to work fine http://paste.openstack.org/show/618171/
13:10:33 vdrok will now look at what's in the nova_api db
13:13:31 figleaf wow, it usually takes me 5 minutes to read the overnight scrollback. Today it was more like 20
13:14:10 figleaf so... anything I can pitch in and help with right now?
13:16:07 mriedem vdrok: yeah i suppose that's why we get here now
13:16:08 mriedem Aug 11 13:07:03 ubuntu nova-compute[11924]: DEBUG nova.scheduler.client.report [None req-b88f01be-3920-4bc0-8708-b96b4f9e8aa7 None None] Updated inventory for 935678ef-67b2-440d-8190-875fb6dea1c6 at generation 3 {{(pid=11924) _update_inventory_attempt /opt/stack/nova/nova/scheduler/client/report.py:652}}
13:16:14 mriedem figleaf: https://etherpad.openstack.org/p/nova-ironic-resource-class-migration
13:16:17 vdrok dtantsur: mriedem what's in nova_api seems to be correct, instance_extra.flavor too, but see this http://paste.openstack.org/show/618173/
13:16:24 figleaf mriedem: yeah, got that open already
13:16:35 vdrok namely, negative free values in compute_nodes
13:16:53 mriedem vdrok: that might be a latent problem?
13:17:01 mriedem i never look at hypervisor-stats, especially for ironic
13:17:03 vdrok mriedem: might be yeah
13:17:17 vdrok I'll try to boot another instance
13:17:26 mriedem vdrok: https://bugs.launchpad.net/nova/+bug/1699947 ?
13:17:27 openstack Launchpad bug 1699947 in OpenStack Compute (nova) "nova hypervisor-stats/hypervisor-show shows wrong resource usage for baremetal node" [Low,Confirmed]
13:17:44 vdrok yup, that is an old one :)
13:18:57 vdrok ok, scheduling seems to work fine too
13:20:51 dtantsur sweet! thanks for testing vdrok :)
13:21:10 dtantsur so, what are the next steps? wait for the CI, check that the warning is gone, then write some unit tests and merge?
13:23:13 dtantsur mriedem: do you think it will still try to delete the allocation?
13:25:23 openstackgerrit Matt Riedemann proposed openstack/nova master: doc: add another up-call caveat for cells v2 for xenapi aggregates https://review.openstack.org/493006
13:25:23 openstackgerrit Matt Riedemann proposed openstack/nova master: doc: add superconductor up-call caveat for cross_az_attach=False https://review.openstack.org/493007
13:25:25 mriedem dansmith: melwitt: ^ superconductor up-call limitations for the docs - i found another one today
13:25:38 mriedem dtantsur: you mean the inventory?
13:25:53 dtantsur mriedem: yes. sorry, tired already :)
13:26:02 dtantsur Friday is not the best day to debug Nova :)
13:26:12 mriedem dtantsur: no because inv_data will not be empty https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L779
13:26:44 dtantsur mriedem: right, so the warning should be gone, no? we won't try to delete it, just update?
13:27:13 mriedem dtantsur: i'd expect to get a 409 response from placement here https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L566
13:27:38 mriedem dtantsur: plus, that doesn't really fix the bug that exists in ocata already, because your fix would only bypass the delete_inventory path iff there is a node.resource_class set
13:27:54 mriedem dtantsur: in other words, i think there are two bugs
13:28:01 dtantsur the 4th problems \o/
13:28:13 dtantsur should I mark it as Related-Bug then?
13:28:16 mriedem the one i reported about the warnings is latent, and exists in ocata
13:28:21 mriedem dtantsur: that would be ok probably
13:28:40 mriedem i.e. in ocata, if i've got a baremetal env, i'm going to see these warnings in n-cpu every 60 seconds for all nodes
13:28:51 mriedem for all *consumed* nodes
13:28:54 mriedem which is annoying
13:29:37 dtantsur mriedem: wait, why? we no longer return an empty inventory for valid nodes. we return an inventory with s/vcpus/vcpus_used/. I think we can backport it to Ocata even
13:30:14 dtantsur jaypipes: hi, you may want to join the party :)
13:30:20 mriedem oh i see
13:30:27 mriedem i hadn't seen the latest patch https://review.openstack.org/#/c/492964/3/nova/virt/ironic/driver.py
13:30:50 dtantsur ah!
13:30:51 jaypipes dtantsur: what kind of party? ;)
13:30:58 dtantsur jaypipes: an ironic party!
13:31:07 dtantsur leakypipes: https://etherpad.openstack.org/p/nova-ironic-resource-class-migration
13:32:05 mriedem dtantsur: yeah that might just work
13:32:15 dtantsur okay, let's wait for Jenkins
13:34:07 mriedem ok so the *_used values for inventory will/should match in placement what we have consumed for allocations on that node provider
13:34:20 mriedem so we shouldn't try to remove any inventory, and thus avoid the 409,
13:34:29 mriedem and create the custom resource class when it's added to the node,
13:34:41 cdent that’s the hoe
13:34:42 mriedem and take the node out of scheduling decisions since inventory == allocation
13:34:43 cdent hope
13:34:49 mriedem who you callin a ho
13:34:50 mriedem ?!
13:34:53 dtantsur LOOOL
13:35:11 cdent hoe for capitalism
13:35:14 dtantsur but yeah, this is the plan
13:42:49 cdent gibi is an evolved tool user. on https://review.openstack.org/#/c/491529/ are you saying you think you found a bug in shelve/unshelve itself, or in the tests?
13:44:59 cdent dtantsur: because it has been shortcutted: tox -epy27 ironic
13:45:12 gibi cdent: I think it is in the shelve/unshelve
13:45:24 gibi cdent: but I'm still busy with the evacuation fix
13:45:26 cdent gibi: go you. your powers are strong.
13:45:51 dtantsur awesome, thanks cdent
13:46:12 gibi cdent: after I pushed the evac fix I can create a shelve/unshelve test
13:46:43 cdent there’s loads of random code floating around
13:49:40 mriedem afk for a bit
14:12:34 openstackgerrit Dmitry Tantsur proposed openstack/nova master: Fix reporting inventory for the Ironic driver https://review.openstack.org/492964
14:12:35 dtantsur cdent, mriedem, bauzas, cleaned up version of my patch ^^^
14:13:19 dtantsur aaaaand the patch using resource classes has passed CI: https://review.openstack.org/#/c/476968/
14:13:26 dtantsur bauzas: you may want to check it for logging ^^^

Earlier   Later