Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-11
11:16:03 openstackgerrit Sylvain Bauza proposed openstack/nova master: Handle addition of new nodes/instances in ironic flavor migration https://review.openstack.org/487954
11:16:17 mriedem dtantsur: i was not sure what to make of http://logs.openstack.org/68/476968/12/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial/02053cf/logs/screen-n-sch.txt.gz#_Aug_09_20_35_41_533621
11:16:27 mriedem dtantsur: when i dug through those logs,
11:16:37 mriedem the ironic driver wasn't populating inventory in placement,
11:16:44 mriedem which would have auto-created the custom resource class
11:17:01 mriedem i'm not sure if we have some chicken and egg issue
11:17:15 dtantsur mriedem: it seems to be that it tries to proceed, and actually fails on http://logs.openstack.org/68/476968/12/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial/02053cf/logs/screen-n-sch.txt.gz#_Aug_09_20_35_41_588272
11:17:21 dtantsur I'm not sure it's expected or not
11:17:37 mriedem dtantsur: i think that's a side effect
11:17:37 dtantsur if it is, then I'm pretty sure we have a chicked and egg situation
11:18:04 mriedem there is a periodic task in the compute service that pulls inventory from ironic https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L775
11:18:13 mriedem ^ includes any custom resource class on the node
11:18:42 mriedem the resource tracker in the compute service calls that from here https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L834
11:19:11 mriedem and if there is a custom resource class in that inventory, it would auto-create it in placement here https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L775
11:19:29 mriedem when i was looking at the logs on that failed job, i never saw _update_inventory get called
11:20:16 mriedem so with my limited understanding of how the ironic driver works, when does the node cache in the driver actually have something show up here? https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L735
11:20:27 vdrok 57-44dc-8b96-4c17182a9a64 have been updated for custom resource class 'baremetal'."
11:20:27 vdrok dtantsur: locally, I see "The flavor extra_specs for Ironic instance 0b8460b1-22
11:20:36 cdent mriedem, bauzas any pending crises that need attention?
11:20:55 mriedem cdent: not really, just trying to sort out https://review.openstack.org/#/c/487954/1
11:20:59 mriedem https://review.openstack.org/#/c/487954/ i mean
11:21:05 vdrok dtantsur: mriedem http://paste.openstack.org/show/618160/
11:21:06 mriedem and this http://logs.openstack.org/68/476968/12/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial/02053cf/logs/screen-n-sch.txt.gz#_Aug_09_20_35_41_533621
11:21:39 vdrok so same thing as you see I guess
11:21:59 openstackgerrit Sylvain Bauza proposed openstack/nova master: DNM: Add more logging + hook devstack https://review.openstack.org/492957
11:22:00 bauzas cdent: nothing really critical
11:22:12 bauzas cdent: ironic flavor migration is hold for RC2
11:22:17 bauzas cdent: and RC1 is on its way
11:22:17 mriedem "Unable to allocate inventory for resource provider 68d57495-2daa-46b3-8e2c-f2f0a19dbaf8: No such resource class CUSTOM_BAREMETAL."
11:22:32 bauzas mriedem: dtantsur: https://review.openstack.org/492957 HTH
11:22:55 dtantsur mriedem: I guess it does it once in 2 minutes. this may be the cause of the problem: we may need to wait for Placement to get it. Do we have any CLI for placement so far?
11:23:18 mriedem dtantsur: there is a series of unmerged osc changes
11:23:26 bauzas (lunch)
11:23:30 mriedem https://blueprints.launchpad.net/nova/+spec/placement-osc-plugin
11:23:50 vdrok dtantsur: don't think there is a problem with waiting, I still see the same error after 4 minutes
11:24:18 dtantsur oh
11:24:35 dtantsur thanks vdrok. this does look like a problem to me now
11:25:44 mriedem this is really weird
11:25:45 mriedem Aug 11 11:18:56 ubuntu nova-compute[7647]: INFO nova.scheduler.client.report [None req-267b8c81-5ea5-4e13-9ea8-02354628c37f None None] Compute node 68d57495-2daa-46b3-8e2c-f2f0a19dbaf8 reported no inventory but previous inventory was detected. Deleting existing inventory records.
11:26:00 mriedem ^ is if the get_inventory method in the driver reported some inventory at one point, but now it doesn't
11:26:58 mriedem was the node disabled?
11:27:10 dtantsur vdrok: ^^^ what was your testing scenario?
11:27:12 mriedem https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L741
11:28:38 vdrok dtantsur: mriedem I had an instance booted yesterday, devstack setup also from yesterday, with no custom non-merged patches. then I updated nova code with this flavor migration change, restarted n-cpu, updated the resource class of the node the instance was on from None to baremetal
11:29:09 dtantsur mriedem: oh. it seems that we don't report inventory for ACTIVE nodes
11:29:19 dtantsur https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L324-L329
11:29:27 mriedem Aug 11 11:18:56 ubuntu nova-compute[7647]: DEBUG nova.virt.ironic.driver [None req-267b8c81-5ea5-4e13-9ea8-02354628c37f None None] The flavor extra_specs for Ironic instance 0b8460b1-22 57-44dc-8b96-4c17182a9a64 have been updated for custom resource class 'baremetal'. {{(pid=7647) _pike_flavor_migration /opt/stack/nova/nova/virt/ironic/driver.py:561}}
11:29:27 mriedem vdrok: ok and we see that get picked up here
11:29:39 vdrok yup
11:29:40 dtantsur which is... wrong?
11:30:19 mriedem dtantsur: ACTIVE?
11:30:22 mriedem https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L176
11:30:26 mriedem that state checking seems fine
11:30:33 mriedem if it's not available, don't say it is
11:30:48 dtantsur states.ACTIVE is not in good state
11:30:58 dtantsur which is the state we have when we have an instance provisioned
11:31:26 dtantsur yeah, but should we return an empty inventory for nodes with an instance? actually, I guess, we should, right
11:31:46 dtantsur but maybe instead of an empty dict we should return all values with zeroes?
11:31:46 mriedem no idea
11:31:49 dtantsur mriedem: ^^^
11:31:59 dtantsur like CUSTOM_FOOBAR=0, VCPU=0, etc?
11:32:29 mriedem we can't report 0 inventory to placement
11:32:30 mriedem https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L752
11:32:38 mriedem min_unit is minimum of 1
11:32:52 dtantsur hmmm
11:33:04 dtantsur then the placement indeed has no idea about the new resource class
11:33:13 dtantsur we never return it from virt/ironic
11:33:17 mriedem https://github.com/openstack/nova/blob/master/nova/api/openstack/placement/handlers/inventory.py#L39
11:33:23 mriedem yeah inventory total has to be at least 1
11:33:30 dtantsur but wait, cannot a hypervisor have 0 free memory, for example?
11:33:33 mriedem dtantsur: is that the chicken and egg?
11:33:41 dtantsur (please pardon my ignorance)
11:34:39 dtantsur but yes, I suspect we cannot create resource classes for such nodes, because we don't report them back ever
11:35:19 mriedem ok so the node has the resource_class set after there is an instance associated and that's how we migrate the flavor extra spec,
11:35:35 mriedem but once the instance is associated, the node is in an ACTIVE state which means we don't report inventory for it?
11:35:40 dtantsur correct
11:35:43 mriedem and thus can't auto-create the newly added resource class
11:36:13 mriedem and that instance <> node is what's already consuming the existing vcpu/memory_mb/disk_gb inventory that we can't delete now
11:36:28 mriedem Aug 11 11:18:56 ubuntu nova-compute[7647]: WARNING nova.scheduler.client.report [None req-267b8c81-5ea5-4e13-9ea8-02354628c37f None None] [req-27dbace7-c527-4f8f-add8-d732047d7382] We c annot delete inventory 'VCPU, MEMORY_MB, DISK_GB' for resource provider 68d57495-2daa-46b3-8e2c-f2f0a19dbaf8 because the inventory is in use.
11:36:29 dtantsur also correct
11:37:58 mriedem heh, yeah, that warning shows up all the time in an ironic ci job run
11:38:00 mriedem http://logs.openstack.org/54/487954/12/check/gate-tempest-dsvm-ironic-ipa-wholedisk-bios-agent_ipmitool-tinyipa-ubuntu-xenial-nv/041c03a/logs/screen-n-cpu.txt.gz#_Aug_09_19_31_21_252127
11:38:35 mriedem i'll open a bug for this since i think it's something we haven't considered, obviously
11:38:43 dtantsur yes please
11:39:06 dtantsur mriedem: can we create resource classes in Placement from the ironic driver each time we encounter a new resource_class?
11:39:16 dtantsur or is it a crazy idea for some reason?
11:39:44 mriedem well,
11:39:56 mriedem i suppose the idea we return 0 inventory in this case is because there is an instance consuming the node,
11:39:59 cdent i seem to have been disconnected briefly so going back in the log, saw some discussion about trying to report 0 inventory for a ironic node with an instance on it
11:40:02 cdent this is bad
11:40:05 mriedem so we don't want to have the scheduler think there is a node available
11:40:12 mriedem for building a new instance
11:40:15 dtantsur right
11:40:18 mriedem as the node is wholly consumed
11:40:22 cdent the inventory should be whatever the capacity is, and then allocations to cover the instance
11:40:26 dtantsur right, and we cannot return VCPU=0
11:42:32 dtantsur cdent: I tend to agree, I'm not sure why we do it
11:43:00 cdent it’s not a question of tend to agree. if you’re doing that, it violates the principles of how placement is supposed to work...
11:43:21 cdent min_unit in the ironic case doesn’t make a lot of sense
11:43:43 cdent but for vms where you don’t want to allow people to slice up the host into lots of tiny pieces, it is meaningful

Earlier   Later