| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-11 | |||
| 11:20:27 | vdrok | dtantsur: locally, I see "The flavor extra_specs for Ironic instance 0b8460b1-22 | |
| 11:20:27 | vdrok | 57-44dc-8b96-4c17182a9a64 have been updated for custom resource class 'baremetal'." | |
| 11:20:36 | cdent | mriedem, bauzas any pending crises that need attention? | |
| 11:20:55 | mriedem | cdent: not really, just trying to sort out https://review.openstack.org/#/c/487954/1 | |
| 11:20:59 | mriedem | https://review.openstack.org/#/c/487954/ i mean | |
| 11:21:05 | vdrok | dtantsur: mriedem http://paste.openstack.org/show/618160/ | |
| 11:21:06 | mriedem | and this http://logs.openstack.org/68/476968/12/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial/02053cf/logs/screen-n-sch.txt.gz#_Aug_09_20_35_41_533621 | |
| 11:21:39 | vdrok | so same thing as you see I guess | |
| 11:21:59 | openstackgerrit | Sylvain Bauza proposed openstack/nova master: DNM: Add more logging + hook devstack https://review.openstack.org/492957 | |
| 11:22:00 | bauzas | cdent: nothing really critical | |
| 11:22:12 | bauzas | cdent: ironic flavor migration is hold for RC2 | |
| 11:22:17 | mriedem | "Unable to allocate inventory for resource provider 68d57495-2daa-46b3-8e2c-f2f0a19dbaf8: No such resource class CUSTOM_BAREMETAL." | |
| 11:22:17 | bauzas | cdent: and RC1 is on its way | |
| 11:22:32 | bauzas | mriedem: dtantsur: https://review.openstack.org/492957 HTH | |
| 11:22:55 | dtantsur | mriedem: I guess it does it once in 2 minutes. this may be the cause of the problem: we may need to wait for Placement to get it. Do we have any CLI for placement so far? | |
| 11:23:18 | mriedem | dtantsur: there is a series of unmerged osc changes | |
| 11:23:26 | bauzas | (lunch) | |
| 11:23:30 | mriedem | https://blueprints.launchpad.net/nova/+spec/placement-osc-plugin | |
| 11:23:50 | vdrok | dtantsur: don't think there is a problem with waiting, I still see the same error after 4 minutes | |
| 11:24:18 | dtantsur | oh | |
| 11:24:35 | dtantsur | thanks vdrok. this does look like a problem to me now | |
| 11:25:44 | mriedem | this is really weird | |
| 11:25:45 | mriedem | Aug 11 11:18:56 ubuntu nova-compute[7647]: INFO nova.scheduler.client.report [None req-267b8c81-5ea5-4e13-9ea8-02354628c37f None None] Compute node 68d57495-2daa-46b3-8e2c-f2f0a19dbaf8 reported no inventory but previous inventory was detected. Deleting existing inventory records. | |
| 11:26:00 | mriedem | ^ is if the get_inventory method in the driver reported some inventory at one point, but now it doesn't | |
| 11:26:58 | mriedem | was the node disabled? | |
| 11:27:10 | dtantsur | vdrok: ^^^ what was your testing scenario? | |
| 11:27:12 | mriedem | https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L741 | |
| 11:28:38 | vdrok | dtantsur: mriedem I had an instance booted yesterday, devstack setup also from yesterday, with no custom non-merged patches. then I updated nova code with this flavor migration change, restarted n-cpu, updated the resource class of the node the instance was on from None to baremetal | |
| 11:29:09 | dtantsur | mriedem: oh. it seems that we don't report inventory for ACTIVE nodes | |
| 11:29:19 | dtantsur | https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L324-L329 | |
| 11:29:27 | mriedem | vdrok: ok and we see that get picked up here | |
| 11:29:27 | mriedem | Aug 11 11:18:56 ubuntu nova-compute[7647]: DEBUG nova.virt.ironic.driver [None req-267b8c81-5ea5-4e13-9ea8-02354628c37f None None] The flavor extra_specs for Ironic instance 0b8460b1-22 57-44dc-8b96-4c17182a9a64 have been updated for custom resource class 'baremetal'. {{(pid=7647) _pike_flavor_migration /opt/stack/nova/nova/virt/ironic/driver.py:561}} | |
| 11:29:39 | vdrok | yup | |
| 11:29:40 | dtantsur | which is... wrong? | |
| 11:30:19 | mriedem | dtantsur: ACTIVE? | |
| 11:30:22 | mriedem | https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L176 | |
| 11:30:26 | mriedem | that state checking seems fine | |
| 11:30:33 | mriedem | if it's not available, don't say it is | |
| 11:30:48 | dtantsur | states.ACTIVE is not in good state | |
| 11:30:58 | dtantsur | which is the state we have when we have an instance provisioned | |
| 11:31:26 | dtantsur | yeah, but should we return an empty inventory for nodes with an instance? actually, I guess, we should, right | |
| 11:31:46 | mriedem | no idea | |
| 11:31:46 | dtantsur | but maybe instead of an empty dict we should return all values with zeroes? | |
| 11:31:49 | dtantsur | mriedem: ^^^ | |
| 11:31:59 | dtantsur | like CUSTOM_FOOBAR=0, VCPU=0, etc? | |
| 11:32:29 | mriedem | we can't report 0 inventory to placement | |
| 11:32:30 | mriedem | https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L752 | |
| 11:32:38 | mriedem | min_unit is minimum of 1 | |
| 11:32:52 | dtantsur | hmmm | |
| 11:33:04 | dtantsur | then the placement indeed has no idea about the new resource class | |
| 11:33:13 | dtantsur | we never return it from virt/ironic | |
| 11:33:17 | mriedem | https://github.com/openstack/nova/blob/master/nova/api/openstack/placement/handlers/inventory.py#L39 | |
| 11:33:23 | mriedem | yeah inventory total has to be at least 1 | |
| 11:33:30 | dtantsur | but wait, cannot a hypervisor have 0 free memory, for example? | |
| 11:33:33 | mriedem | dtantsur: is that the chicken and egg? | |
| 11:33:41 | dtantsur | (please pardon my ignorance) | |
| 11:34:39 | dtantsur | but yes, I suspect we cannot create resource classes for such nodes, because we don't report them back ever | |
| 11:35:19 | mriedem | ok so the node has the resource_class set after there is an instance associated and that's how we migrate the flavor extra spec, | |
| 11:35:35 | mriedem | but once the instance is associated, the node is in an ACTIVE state which means we don't report inventory for it? | |
| 11:35:40 | dtantsur | correct | |
| 11:35:43 | mriedem | and thus can't auto-create the newly added resource class | |
| 11:36:13 | mriedem | and that instance <> node is what's already consuming the existing vcpu/memory_mb/disk_gb inventory that we can't delete now | |
| 11:36:28 | mriedem | Aug 11 11:18:56 ubuntu nova-compute[7647]: WARNING nova.scheduler.client.report [None req-267b8c81-5ea5-4e13-9ea8-02354628c37f None None] [req-27dbace7-c527-4f8f-add8-d732047d7382] We c annot delete inventory 'VCPU, MEMORY_MB, DISK_GB' for resource provider 68d57495-2daa-46b3-8e2c-f2f0a19dbaf8 because the inventory is in use. | |
| 11:36:29 | dtantsur | also correct | |
| 11:37:58 | mriedem | heh, yeah, that warning shows up all the time in an ironic ci job run | |
| 11:38:00 | mriedem | http://logs.openstack.org/54/487954/12/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial-nv/041c03a/logs/screen-n-cpu.txt.gz#_Aug_09_19_31_21_252127 | |
| 11:38:35 | mriedem | i'll open a bug for this since i think it's something we haven't considered, obviously | |
| 11:38:43 | dtantsur | yes please | |
| 11:39:06 | dtantsur | mriedem: can we create resource classes in Placement from the ironic driver each time we encounter a new resource_class? | |
| 11:39:16 | dtantsur | or is it a crazy idea for some reason? | |
| 11:39:44 | mriedem | well, | |
| 11:39:56 | mriedem | i suppose the idea we return 0 inventory in this case is because there is an instance consuming the node, | |
| 11:39:59 | cdent | i seem to have been disconnected briefly so going back in the log, saw some discussion about trying to report 0 inventory for a ironic node with an instance on it | |
| 11:40:02 | cdent | this is bad | |
| 11:40:05 | mriedem | so we don't want to have the scheduler think there is a node available | |
| 11:40:12 | mriedem | for building a new instance | |
| 11:40:15 | dtantsur | right | |
| 11:40:18 | mriedem | as the node is wholly consumed | |
| 11:40:22 | cdent | the inventory should be whatever the capacity is, and then allocations to cover the instance | |
| 11:40:26 | dtantsur | right, and we cannot return VCPU=0 | |
| 11:42:32 | dtantsur | cdent: I tend to agree, I'm not sure why we do it | |
| 11:43:00 | cdent | it’s not a question of tend to agree. if you’re doing that, it violates the principles of how placement is supposed to work... | |
| 11:43:21 | cdent | min_unit in the ironic case doesn’t make a lot of sense | |
| 11:43:43 | cdent | but for vms where you don’t want to allow people to slice up the host into lots of tiny pieces, it is meaningful | |
| 11:44:21 | dtantsur | cdent: I used "tend to agree" to designate that I do not know Nova well enough, not to question your findings | |
| 11:44:53 | cdent | I didn’t think you were questioning, I was just reinforcing the point: sounds like weird stuff afoot | |
| 11:45:20 | dtantsur | ok | |
| 11:45:36 | dtantsur | so | |
| 11:45:50 | dtantsur | should we just fix it to remove min_unit and always return the correct inventory? | |
| 11:45:58 | dtantsur | s/correct/complete/ | |
| 11:46:23 | dtantsur | oh, I think I know why it was done | |
| 11:46:37 | cdent | one sec | |
| 11:47:08 | vdrok | so _refresh_cache is called in the driver.get_available_nodes from compute manager's update_available_resource periodic, and we migrate the flavor there. then we call the update_available_resource, which calls get_available_resource. we still include the resoruce_class in the return dict in https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L335. | |
| 11:47:19 | dtantsur | cdent: this is mimicking the old behavior with bare metal nodes. if we always report the complete inventory, e.g. 2Gi of RAM. and there is an instance with 1Gi of RAM. the Placement will think that 1Gi of RAM is still free | |
| 11:47:30 | vdrok | dtantsur: so do you propose to call the placement to create resource class right during the update_available_resource? | |
| 11:47:47 | dtantsur | vdrok: something like that.. but let's figure out the inventory problem first | |
| 11:48:03 | mriedem | https://bugs.launchpad.net/nova/+bug/1710141 | |
| 11:48:04 | openstack | Launchpad bug 1710141 in OpenStack Compute (nova) "Continual warnings in n-cpu logs about being unable to delete inventory for an ironic node with an instance on it" [High,Triaged] | |
| 11:48:09 | mriedem | cdent: dtantsur: vdrok: ^ | |
| 11:48:14 | dtantsur | thanks mriedem | |