Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-11
11:57:19 mriedem when i search for things in the nova docs now, they search all of docs.o.o
11:57:20 cdent dtantsur: the issue is that if there is already inventory in use that is based on vcpus, then you won’t be able to change the inventory
11:57:24 mriedem not just the nova docs
11:57:41 dtantsur cdent: why should I?
11:58:16 dtantsur okay, we're trying to solve two problems at the same time:
11:58:16 cdent as I read that code you are trying to replace existin inventory, but maybe I’m not understanding?
11:58:24 mriedem boolman: no i'm talking about this https://github.com/openstack/nova/blob/master/nova/conf/libvirt.py#L631
11:58:30 dtantsur 1. broken ironic scheduling with custom resource classes
11:58:33 vdrok dtantsur: you mean something like http://paste.openstack.org/show/618161/
11:58:34 vdrok ?
11:58:45 dtantsur 2. warning on trying to delete the inventory, because we stop reporting it correctly
11:58:54 dtantsur I'm trying to fix #1 first, as it's a hard blocker for this work
11:58:56 mriedem https://github.com/openstack/nova/blob/master/nova/virt/libvirt/imagebackend.py#L788
11:59:16 dtantsur vdrok: yes, see the DNM patch I posted above
11:59:42 cdent okay, let’s dismiss #2 entirely for a moment
11:59:45 boolman mriedem: ok thanks, will try that
11:59:54 vdrok dtantsur: ah, right, /me is slow :)
12:00:15 dtantsur cdent: so, I think we all agree that our get_inventory is incorrect. We cannot fix it at once, because of the nature of bare metal nodes. we can fix reporting the custom resource class, and just wait for VCPU handling to be removed in Queens completely. Does it make more sense?
12:00:26 cdent dtantsur: can’t you point me at some irc or test logs where the #1 problem is explained or demonstrated?
12:01:09 mriedem cdent: that's in https://bugs.launchpad.net/nova/+bug/1710141
12:01:10 openstack Launchpad bug 1710141 in OpenStack Compute (nova) "Continual warnings in n-cpu logs about being unable to delete inventory for an ironic node with an instance on it" [High,Triaged]
12:01:12 cdent dtantsur: your dnm code is create an entire new inventory, with just the resource class set
12:01:15 cdent thanks mriedem
12:01:28 dtantsur I think this bug is about #2, not #1
12:01:34 mriedem they are related
12:02:04 mriedem cdent: dtantsur has a patch up in ironic to set a custom resource class on the node and then create an instance on that node, but that's failing here http://logs.openstack.org/68/476968/12/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial/02053cf/logs/screen-n-sch.txt.gz#_Aug_09_20_35_41_533621
12:02:17 mriedem during scheduling, because the custom resource class doesn't exist, and nova never creates it
12:02:26 mriedem nova never gets inventory off the node
12:02:39 mriedem dtantsur: that does confuse me though, are we racing?
12:02:56 mriedem wouldn't nova-compute need to report the node before the scheduler could use it anyway?
12:03:14 dtantsur mriedem: this is during a migration from resource_class=None to resource_class=something
12:03:24 dtantsur so the resource class was never reported
12:03:37 mriedem dtantsur: yeah but in that log ^ there was never an instance created yet i don't tihnk
12:03:45 dtantsur ok, so we have 3 problems :)
12:03:53 cdent :)
12:03:58 mriedem basically, when does the driver report that there are available nodes?
12:04:07 mriedem b/c once it does, then we pull inventory from those nodes and put that in placement
12:04:13 mriedem which the scheduler will use to create an instance
12:04:17 dtantsur it polls ironic once in IIRC 2 minutes
12:04:22 dtantsur so there is some space for a race indeed
12:04:57 mriedem but how would this not fail all of the ironic ci jobs that use nova to create instances?
12:05:01 aarefiev dtantsur: but we report empty inventory, right https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L741
12:05:02 mriedem regardless of the resource class stuff
12:05:17 dtantsur mriedem: we used to wait for 'nova hypervisor-stats' to show available CPUS > 0
12:05:23 mriedem dtantsur: and you're not now?
12:05:48 dtantsur mriedem: that patch tries not reporting VCPU any more
12:06:00 dtantsur maybe it's a bit premature
12:06:09 mriedem ok, so we don't wait for the compute node to be registered in nova-compute,
12:06:14 mriedem and try to create the instance,
12:06:19 mriedem and since we didn't wait, we don't have the resource class created yet
12:06:24 mriedem and we NoValidHost
12:06:26 mriedem like a champ
12:06:43 dtantsur right, maybe a wait loop until we get something in placement (hence my question about its CLI) solves it
12:06:48 sdague mriedem: yes, not just all of docs.o.o, also ask.o.o
12:07:01 sdague mriedem: that's definitely an issue
12:07:15 sdague mriedem: going to have to be brought up at PTG I think
12:07:27 cdent dtantsur: placement is so easy to curl that no one has bothered yet
12:07:45 mriedem sdague: complaining in -doc
12:08:01 mriedem cdent: there have been unmerged patches
12:08:08 mriedem cdent: https://review.openstack.org/#/q/project:openstack/osc-placement
12:08:10 cdent mriedem: yes, I know
12:08:14 mriedem ok
12:08:35 cdent i’m one of the few reviewers on those patches, and mentioned them for several months on the rp update weekly messages and finally stopped when no one was reviewing
12:08:43 mriedem :(
12:08:43 cdent because I assumed nobody cared
12:08:53 cdent :( is right
12:09:01 mriedem once we want to start integrating them into CI, people will care
12:09:03 mriedem well, dev people
12:10:13 cdent dtantsur, mriedem: so do you think you’ve gotten past at least a first hurdle with the conceptual wait loop? If so, once that’s cleared out, I’d like eventually to come back to this issue of reporting or not report inventory for nodes that have instances on them
12:10:22 dtantsur folks, I've dumped my/our findings on https://etherpad.openstack.org/p/nova-ironic-resource-class-migration
12:10:30 dtantsur I cannot keep so much in my head :)
12:10:35 cdent good idea
12:11:58 cdent if we merge this those silly html error responses (in the pastes there) will go away: https://review.openstack.org/#/c/489772/
12:13:38 mriedem so like this http://logs.openstack.org/72/489772/2/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial-nv/655a0b3/logs/screen-n-cpu.txt.gz?level=TRACE#_Aug_02_12_12_16_584946
12:14:56 cdent mriedem: yeah, no line feeds...
12:16:29 mriedem like this http://logs.openstack.org/85/490085/7/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial-nv/5835515/logs/screen-n-cpu.txt.gz?level=TRACE#_Aug_03_15_47_53_683694
12:16:42 cdent right
12:23:39 cdent dtantsur: so currently plan is is wait and see how https://review.openstack.org/#/c/476968/ turns out?
12:24:27 mriedem sdague: http://lists.openstack.org/pipermail/openstack-dev/2017-August/121042.html
12:24:42 dtantsur cdent: this is the plan for #1. for #2 and #3, I can try fixing get_inventory indeed.
12:25:14 cdent dtantsur: let me know if there’s something I can help with
12:25:31 dtantsur sure, thanks!
12:29:26 openstackgerrit OpenStack Release Bot proposed openstack/nova master: Update reno for stable/pike https://review.openstack.org/492982
12:32:55 openstackgerrit Dmitry Tantsur proposed openstack/nova master: DNM PoC for fixing ironic with resource classes https://review.openstack.org/492964
12:32:57 dtantsur cdent: something like ^^^?
12:33:55 dtantsur mriedem: ^^
12:34:52 boolman mriedem: ok I got it to work now, thanks
12:37:42 dtantsur vdrok: mind testing again with my patch above?
12:38:02 cdent dtantsur: yes. Was there some additional thing to do to make sure that get_inventory gets called often enough? I’m guessing (giving the log messages) that that’s not a problem?
12:38:10 vdrok dtantsur: ok, will do
12:38:33 dtantsur cdent: I'm not sure, let's see how it looks for vdrok
12:38:38 vdrok yeah, after instance deletion we'll have some time window when the resources reported by placement will be incorrect
12:38:47 cdent ✔
12:40:55 sdague mriedem: I'll see if I can hack around it
12:41:24 dtantsur once we get custom resource classes to work with ironic, we can ask operators to upgrade. then they won't see issues with VCPU reporting..
12:44:04 cdent biab
12:46:41 bauzas mriedem: dtantsur: could you please tl;dr the issues with ironic ?
12:46:50 bauzas and how I could help ?
12:47:11 bauzas tons of channel logs :)
12:48:26 dtantsur bauzas: this is the tl;dr https://etherpad.openstack.org/p/nova-ironic-resource-class-migration

Earlier   Later