| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-11 | |||
| 11:48:14 | dtantsur | thanks mriedem | |
| 11:49:19 | dtantsur | anyway, cdent, mriedem, wdyt about returning the custom resource class in the inventory *always*, even for occupied nodes? | |
| 11:49:21 | vdrok | dtantsur: I think the reason of having min_unit=0 is because if the resource provider can provide 0 of some resource, it;s not really that resource's provider :) | |
| 11:49:30 | vdrok | err, min_unit=1 | |
| 11:49:58 | vdrok | if that's what you're talking about | |
| 11:50:24 | cdent | dtantsur: are you talking about the get_inventory call in the virt driver? | |
| 11:50:40 | dtantsur | cdent: yes. maybe I should make a DNM patch showing it.. | |
| 11:50:59 | cdent | if the physical hardware hasn’t changed, that should always return the same thing, without regard to presence of an instance | |
| 11:51:14 | cdent | so I think the answer to your question is "yes" | |
| 11:51:40 | mriedem | as cdent pointed out, it seems we should be reporting the inventory regardless of what's allocated on that provider, and let the allocation consume the inventory so the scheduler will ignore it | |
| 11:51:55 | mriedem | i.e. this node has 1 VCPU and that 1 VCPU is consumed, so it's not eligible for building another instance | |
| 11:52:00 | boolman | hi peeps, Can I disable local disk on hypervisors? eg: openstack server create --image xenial --security-group default --key-name emil --network backend --flavor smallish demo -- currently this creates the instance on local disk on the hypervisors | |
| 11:52:21 | mriedem | boolman: you'd have to use a volume | |
| 11:52:24 | boolman | since I'm using ceph rbd storage i want to force that | |
| 11:52:28 | openstackgerrit | Dmitry Tantsur proposed openstack/nova master: DNM PoC for fixing ironic with resource classes https://review.openstack.org/492964 | |
| 11:52:29 | dtantsur | cdent, mriedem, something like ^^^ | |
| 11:52:46 | mriedem | boolman: boot from volume that deletes on termination | |
| 11:52:49 | openstackgerrit | Chris Dent proposed openstack/nova master: Optional separate database for placement API https://review.openstack.org/362766 | |
| 11:53:35 | boolman | mriedem: so I can't actually disable local disk? the user have to create a volume to use when creating the instance? | |
| 11:54:01 | mriedem | dtantsur: i'm not sure that will work, i'd expect the PUT /resource_providers/uuid/inventories to fail to remove the inventory for the vcpu/memory_mb/disk_gb because it's already being used | |
| 11:54:09 | mriedem | it would auto-create the custom resource class though, | |
| 11:54:12 | mriedem | if that's your aim | |
| 11:54:38 | dtantsur | mriedem: it already fails at removing the inventory. I'm trying to fix auto-creation for now | |
| 11:54:46 | mriedem | boolman: you could use the rbd imagebackend on the compute so your computes are shared a ceph pool of disk | |
| 11:54:51 | mriedem | if you don't want local being used | |
| 11:54:57 | mriedem | *sharing | |
| 11:55:31 | mriedem | dtantsur: yeah, this is just probably not the way to do this i don't think | |
| 11:55:42 | boolman | mriedem: you mean by modifying the pool in virsh? | |
| 11:55:48 | mriedem | it's super tightly coupled to knowing exactly how the inventory is used by the RT and the report client | |
| 11:56:02 | mriedem | boolman: see http://lists.openstack.org/pipermail/openstack-dev/2017-May/117012.html | |
| 11:56:47 | dtantsur | mriedem: but isn't it the correct thing to do? I mean, always return the inventory of this custom resource class? | |
| 11:57:03 | dtantsur | (given that we will remove the hacks around VCPU and friends in Queens) | |
| 11:57:10 | mriedem | sdague: docs migration annoyance of the day, | |
| 11:57:19 | mriedem | when i search for things in the nova docs now, they search all of docs.o.o | |
| 11:57:20 | cdent | dtantsur: the issue is that if there is already inventory in use that is based on vcpus, then you won’t be able to change the inventory | |
| 11:57:24 | mriedem | not just the nova docs | |
| 11:57:41 | dtantsur | cdent: why should I? | |
| 11:58:16 | dtantsur | okay, we're trying to solve two problems at the same time: | |
| 11:58:16 | cdent | as I read that code you are trying to replace existin inventory, but maybe I’m not understanding? | |
| 11:58:24 | mriedem | boolman: no i'm talking about this https://github.com/openstack/nova/blob/master/nova/conf/libvirt.py#L631 | |
| 11:58:30 | dtantsur | 1. broken ironic scheduling with custom resource classes | |
| 11:58:33 | vdrok | dtantsur: you mean something like http://paste.openstack.org/show/618161/ | |
| 11:58:34 | vdrok | ? | |
| 11:58:45 | dtantsur | 2. warning on trying to delete the inventory, because we stop reporting it correctly | |
| 11:58:54 | dtantsur | I'm trying to fix #1 first, as it's a hard blocker for this work | |
| 11:58:56 | mriedem | https://github.com/openstack/nova/blob/master/nova/virt/libvirt/imagebackend.py#L788 | |
| 11:59:16 | dtantsur | vdrok: yes, see the DNM patch I posted above | |
| 11:59:42 | cdent | okay, let’s dismiss #2 entirely for a moment | |
| 11:59:45 | boolman | mriedem: ok thanks, will try that | |
| 11:59:54 | vdrok | dtantsur: ah, right, /me is slow :) | |
| 12:00:15 | dtantsur | cdent: so, I think we all agree that our get_inventory is incorrect. We cannot fix it at once, because of the nature of bare metal nodes. we can fix reporting the custom resource class, and just wait for VCPU handling to be removed in Queens completely. Does it make more sense? | |
| 12:00:26 | cdent | dtantsur: can’t you point me at some irc or test logs where the #1 problem is explained or demonstrated? | |
| 12:01:09 | mriedem | cdent: that's in https://bugs.launchpad.net/nova/+bug/1710141 | |
| 12:01:10 | openstack | Launchpad bug 1710141 in OpenStack Compute (nova) "Continual warnings in n-cpu logs about being unable to delete inventory for an ironic node with an instance on it" [High,Triaged] | |
| 12:01:12 | cdent | dtantsur: your dnm code is create an entire new inventory, with just the resource class set | |
| 12:01:15 | cdent | thanks mriedem | |
| 12:01:28 | dtantsur | I think this bug is about #2, not #1 | |
| 12:01:34 | mriedem | they are related | |
| 12:02:04 | mriedem | cdent: dtantsur has a patch up in ironic to set a custom resource class on the node and then create an instance on that node, but that's failing here http://logs.openstack.org/68/476968/12/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial/02053cf/logs/screen-n-sch.txt.gz#_Aug_09_20_35_41_533621 | |
| 12:02:17 | mriedem | during scheduling, because the custom resource class doesn't exist, and nova never creates it | |
| 12:02:26 | mriedem | nova never gets inventory off the node | |
| 12:02:39 | mriedem | dtantsur: that does confuse me though, are we racing? | |
| 12:02:56 | mriedem | wouldn't nova-compute need to report the node before the scheduler could use it anyway? | |
| 12:03:14 | dtantsur | mriedem: this is during a migration from resource_class=None to resource_class=something | |
| 12:03:24 | dtantsur | so the resource class was never reported | |
| 12:03:37 | mriedem | dtantsur: yeah but in that log ^ there was never an instance created yet i don't tihnk | |
| 12:03:45 | dtantsur | ok, so we have 3 problems :) | |
| 12:03:53 | cdent | :) | |
| 12:03:58 | mriedem | basically, when does the driver report that there are available nodes? | |
| 12:04:07 | mriedem | b/c once it does, then we pull inventory from those nodes and put that in placement | |
| 12:04:13 | mriedem | which the scheduler will use to create an instance | |
| 12:04:17 | dtantsur | it polls ironic once in IIRC 2 minutes | |
| 12:04:22 | dtantsur | so there is some space for a race indeed | |
| 12:04:57 | mriedem | but how would this not fail all of the ironic ci jobs that use nova to create instances? | |
| 12:05:01 | aarefiev | dtantsur: but we report empty inventory, right https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L741 | |
| 12:05:02 | mriedem | regardless of the resource class stuff | |
| 12:05:17 | dtantsur | mriedem: we used to wait for 'nova hypervisor-stats' to show available CPUS > 0 | |
| 12:05:23 | mriedem | dtantsur: and you're not now? | |
| 12:05:48 | dtantsur | mriedem: that patch tries not reporting VCPU any more | |
| 12:06:00 | dtantsur | maybe it's a bit premature | |
| 12:06:09 | mriedem | ok, so we don't wait for the compute node to be registered in nova-compute, | |
| 12:06:14 | mriedem | and try to create the instance, | |
| 12:06:19 | mriedem | and since we didn't wait, we don't have the resource class created yet | |
| 12:06:24 | mriedem | and we NoValidHost | |
| 12:06:26 | mriedem | like a champ | |
| 12:06:43 | dtantsur | right, maybe a wait loop until we get something in placement (hence my question about its CLI) solves it | |
| 12:06:48 | sdague | mriedem: yes, not just all of docs.o.o, also ask.o.o | |
| 12:07:01 | sdague | mriedem: that's definitely an issue | |
| 12:07:15 | sdague | mriedem: going to have to be brought up at PTG I think | |
| 12:07:27 | cdent | dtantsur: placement is so easy to curl that no one has bothered yet | |
| 12:07:45 | mriedem | sdague: complaining in -doc | |
| 12:08:01 | mriedem | cdent: there have been unmerged patches | |
| 12:08:08 | mriedem | cdent: https://review.openstack.org/#/q/project:openstack/osc-placement | |
| 12:08:10 | cdent | mriedem: yes, I know | |
| 12:08:14 | mriedem | ok | |
| 12:08:35 | cdent | i’m one of the few reviewers on those patches, and mentioned them for several months on the rp update weekly messages and finally stopped when no one was reviewing | |
| 12:08:43 | mriedem | :( | |
| 12:08:43 | cdent | because I assumed nobody cared | |
| 12:08:53 | cdent | :( is right | |
| 12:09:01 | mriedem | once we want to start integrating them into CI, people will care | |