| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-11 | |||
| 11:58:16 | dtantsur | okay, we're trying to solve two problems at the same time: | |
| 11:58:16 | cdent | as I read that code you are trying to replace existin inventory, but maybe I’m not understanding? | |
| 11:58:24 | mriedem | boolman: no i'm talking about this https://github.com/openstack/nova/blob/master/nova/conf/libvirt.py#L631 | |
| 11:58:30 | dtantsur | 1. broken ironic scheduling with custom resource classes | |
| 11:58:33 | vdrok | dtantsur: you mean something like http://paste.openstack.org/show/618161/ | |
| 11:58:34 | vdrok | ? | |
| 11:58:45 | dtantsur | 2. warning on trying to delete the inventory, because we stop reporting it correctly | |
| 11:58:54 | dtantsur | I'm trying to fix #1 first, as it's a hard blocker for this work | |
| 11:58:56 | mriedem | https://github.com/openstack/nova/blob/master/nova/virt/libvirt/imagebackend.py#L788 | |
| 11:59:16 | dtantsur | vdrok: yes, see the DNM patch I posted above | |
| 11:59:42 | cdent | okay, let’s dismiss #2 entirely for a moment | |
| 11:59:45 | boolman | mriedem: ok thanks, will try that | |
| 11:59:54 | vdrok | dtantsur: ah, right, /me is slow :) | |
| 12:00:15 | dtantsur | cdent: so, I think we all agree that our get_inventory is incorrect. We cannot fix it at once, because of the nature of bare metal nodes. we can fix reporting the custom resource class, and just wait for VCPU handling to be removed in Queens completely. Does it make more sense? | |
| 12:00:26 | cdent | dtantsur: can’t you point me at some irc or test logs where the #1 problem is explained or demonstrated? | |
| 12:01:09 | mriedem | cdent: that's in https://bugs.launchpad.net/nova/+bug/1710141 | |
| 12:01:10 | openstack | Launchpad bug 1710141 in OpenStack Compute (nova) "Continual warnings in n-cpu logs about being unable to delete inventory for an ironic node with an instance on it" [High,Triaged] | |
| 12:01:12 | cdent | dtantsur: your dnm code is create an entire new inventory, with just the resource class set | |
| 12:01:15 | cdent | thanks mriedem | |
| 12:01:28 | dtantsur | I think this bug is about #2, not #1 | |
| 12:01:34 | mriedem | they are related | |
| 12:02:04 | mriedem | cdent: dtantsur has a patch up in ironic to set a custom resource class on the node and then create an instance on that node, but that's failing here http://logs.openstack.org/68/476968/12/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial/02053cf/logs/screen-n-sch.txt.gz#_Aug_09_20_35_41_533621 | |
| 12:02:17 | mriedem | during scheduling, because the custom resource class doesn't exist, and nova never creates it | |
| 12:02:26 | mriedem | nova never gets inventory off the node | |
| 12:02:39 | mriedem | dtantsur: that does confuse me though, are we racing? | |
| 12:02:56 | mriedem | wouldn't nova-compute need to report the node before the scheduler could use it anyway? | |
| 12:03:14 | dtantsur | mriedem: this is during a migration from resource_class=None to resource_class=something | |
| 12:03:24 | dtantsur | so the resource class was never reported | |
| 12:03:37 | mriedem | dtantsur: yeah but in that log ^ there was never an instance created yet i don't tihnk | |
| 12:03:45 | dtantsur | ok, so we have 3 problems :) | |
| 12:03:53 | cdent | :) | |
| 12:03:58 | mriedem | basically, when does the driver report that there are available nodes? | |
| 12:04:07 | mriedem | b/c once it does, then we pull inventory from those nodes and put that in placement | |
| 12:04:13 | mriedem | which the scheduler will use to create an instance | |
| 12:04:17 | dtantsur | it polls ironic once in IIRC 2 minutes | |
| 12:04:22 | dtantsur | so there is some space for a race indeed | |
| 12:04:57 | mriedem | but how would this not fail all of the ironic ci jobs that use nova to create instances? | |
| 12:05:01 | aarefiev | dtantsur: but we report empty inventory, right https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L741 | |
| 12:05:02 | mriedem | regardless of the resource class stuff | |
| 12:05:17 | dtantsur | mriedem: we used to wait for 'nova hypervisor-stats' to show available CPUS > 0 | |
| 12:05:23 | mriedem | dtantsur: and you're not now? | |
| 12:05:48 | dtantsur | mriedem: that patch tries not reporting VCPU any more | |
| 12:06:00 | dtantsur | maybe it's a bit premature | |
| 12:06:09 | mriedem | ok, so we don't wait for the compute node to be registered in nova-compute, | |
| 12:06:14 | mriedem | and try to create the instance, | |
| 12:06:19 | mriedem | and since we didn't wait, we don't have the resource class created yet | |
| 12:06:24 | mriedem | and we NoValidHost | |
| 12:06:26 | mriedem | like a champ | |
| 12:06:43 | dtantsur | right, maybe a wait loop until we get something in placement (hence my question about its CLI) solves it | |
| 12:06:48 | sdague | mriedem: yes, not just all of docs.o.o, also ask.o.o | |
| 12:07:01 | sdague | mriedem: that's definitely an issue | |
| 12:07:15 | sdague | mriedem: going to have to be brought up at PTG I think | |
| 12:07:27 | cdent | dtantsur: placement is so easy to curl that no one has bothered yet | |
| 12:07:45 | mriedem | sdague: complaining in -doc | |
| 12:08:01 | mriedem | cdent: there have been unmerged patches | |
| 12:08:08 | mriedem | cdent: https://review.openstack.org/#/q/project:openstack/osc-placement | |
| 12:08:10 | cdent | mriedem: yes, I know | |
| 12:08:14 | mriedem | ok | |
| 12:08:35 | cdent | i’m one of the few reviewers on those patches, and mentioned them for several months on the rp update weekly messages and finally stopped when no one was reviewing | |
| 12:08:43 | mriedem | :( | |
| 12:08:43 | cdent | because I assumed nobody cared | |
| 12:08:53 | cdent | :( is right | |
| 12:09:01 | mriedem | once we want to start integrating them into CI, people will care | |
| 12:09:03 | mriedem | well, dev people | |
| 12:10:13 | cdent | dtantsur, mriedem: so do you think you’ve gotten past at least a first hurdle with the conceptual wait loop? If so, once that’s cleared out, I’d like eventually to come back to this issue of reporting or not report inventory for nodes that have instances on them | |
| 12:10:22 | dtantsur | folks, I've dumped my/our findings on https://etherpad.openstack.org/p/nova-ironic-resource-class-migration | |
| 12:10:30 | dtantsur | I cannot keep so much in my head :) | |
| 12:10:35 | cdent | good idea | |
| 12:11:58 | cdent | if we merge this those silly html error responses (in the pastes there) will go away: https://review.openstack.org/#/c/489772/ | |
| 12:13:38 | mriedem | so like this http://logs.openstack.org/72/489772/2/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial-nv/655a0b3/logs/screen-n-cpu.txt.gz?level=TRACE#_Aug_02_12_12_16_584946 | |
| 12:14:56 | cdent | mriedem: yeah, no line feeds... | |
| 12:16:29 | mriedem | like this http://logs.openstack.org/85/490085/7/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial-nv/5835515/logs/screen-n-cpu.txt.gz?level=TRACE#_Aug_03_15_47_53_683694 | |
| 12:16:42 | cdent | right | |
| 12:23:39 | cdent | dtantsur: so currently plan is is wait and see how https://review.openstack.org/#/c/476968/ turns out? | |
| 12:24:27 | mriedem | sdague: http://lists.openstack.org/pipermail/openstack-dev/2017-August/121042.html | |
| 12:24:42 | dtantsur | cdent: this is the plan for #1. for #2 and #3, I can try fixing get_inventory indeed. | |
| 12:25:14 | cdent | dtantsur: let me know if there’s something I can help with | |
| 12:25:31 | dtantsur | sure, thanks! | |
| 12:29:26 | openstackgerrit | OpenStack Release Bot proposed openstack/nova master: Update reno for stable/pike https://review.openstack.org/492982 | |
| 12:32:55 | openstackgerrit | Dmitry Tantsur proposed openstack/nova master: DNM PoC for fixing ironic with resource classes https://review.openstack.org/492964 | |
| 12:32:57 | dtantsur | cdent: something like ^^^? | |
| 12:33:55 | dtantsur | mriedem: ^^ | |
| 12:34:52 | boolman | mriedem: ok I got it to work now, thanks | |
| 12:37:42 | dtantsur | vdrok: mind testing again with my patch above? | |
| 12:38:02 | cdent | dtantsur: yes. Was there some additional thing to do to make sure that get_inventory gets called often enough? I’m guessing (giving the log messages) that that’s not a problem? | |
| 12:38:10 | vdrok | dtantsur: ok, will do | |
| 12:38:33 | dtantsur | cdent: I'm not sure, let's see how it looks for vdrok | |
| 12:38:38 | vdrok | yeah, after instance deletion we'll have some time window when the resources reported by placement will be incorrect | |
| 12:38:47 | cdent | ✔ | |
| 12:40:55 | sdague | mriedem: I'll see if I can hack around it | |
| 12:41:24 | dtantsur | once we get custom resource classes to work with ironic, we can ask operators to upgrade. then they won't see issues with VCPU reporting.. | |
| 12:44:04 | cdent | biab | |
| 12:46:41 | bauzas | mriedem: dtantsur: could you please tl;dr the issues with ironic ? | |
| 12:46:50 | bauzas | and how I could help ? | |
| 12:47:11 | bauzas | tons of channel logs :) | |
| 12:48:26 | dtantsur | bauzas: this is the tl;dr https://etherpad.openstack.org/p/nova-ironic-resource-class-migration | |
| 12:48:47 | bauzas | excellent, thanks | |
| 12:54:06 | vdrok | dtantsur: see comment | |
| 12:54:15 | vdrok | right now requests to placement fail | |
| 12:54:21 | vdrok | because of max_unit=0 | |