Earlier  
Posted Nick Remark
#openstack-nova - 2018-08-29
14:01:03 mriedem zigo: yes
14:01:17 mriedem zigo: naichuans: if you haven't already, it would be great to have a nova bug for tracking this
14:01:30 cdent (while still keeping a record of the "actual" inventory)
14:01:32 mriedem yeah idk, maybe jaypipes remmebers
14:02:09 mriedem so https://github.com/openstack/nova/blob/6522ea3ecfe99cca3fb33258b11e5a1f34e6e8f0/nova/compute/resource_tracker.py#L84 in the RT is what is supposed to set the default allocation ratio for each inventory on the compute node provider
14:02:24 zigo mriedem: cdent: What is max_unit supposed to represent? Isn't it total * allocation_ratio?
14:02:24 mriedem assuming the virt driver doesn't provide an allocation ratio override, which neither libvirt nor xenapi do
14:02:43 jaypipes mriedem: sorry, reading back was on a call
14:02:55 cdent zigo: no, it is an expression of the largest amount any individual resource allocation can allocate
14:02:57 jaypipes zigo: max_unit is not total * allocation_ratio, no
14:03:07 cdent so in the case of something like disk or vcpu it is generally the max physical amount
14:03:24 cdent as you don't want a single instance to have more than is physically available, regardless of allocation ratio
14:03:37 zigo Ah, great, meaning I can have hosts with up to 2 billion VCUPs ! :)
14:04:13 zigo I can start writing another bug, ok.
14:04:44 cdent total: the real physical amount, allocation ratio: multiplier for over or under commit, min_unit: smallest individual request, max_unit: largest individual request, reserved: what the system is using for itself
14:06:24 mriedem i wonder if there is a bug in ProviderTree.update_inventory where it's not updating the allocation_ratio when we create/update the provider in placement
14:07:23 mriedem zigo: you're testing with rocky right?
14:10:55 openstack Launchpad bug 1789654 in OpenStack Compute (nova) "placement allocation_ratio initialized with 0.0" [Undecided,New]
14:10:55 zigo https://bugs.launchpad.net/nova/+bug/1789654
14:10:59 zigo mriedem: Correct !
14:11:15 zigo mriedem: Fresh packages just built this and last week in a record time.
14:11:24 zigo It took me less than 2 weeks, this time ... :P
14:11:36 openstackgerrit Merged openstack/osc-placement master: Update reno for stable/rocky https://review.openstack.org/586115
14:13:34 mriedem so i thought we used to log something when inventory changed on a provider...
14:13:36 mriedem not seeing that
14:13:51 mriedem from one of the failed xenserver ci logs, the RP is created here http://logs.openstack.org/41/590041/17/check/tempest-full/b3f9ddd/controller/logs/screen-n-cpu.txt.gz#_Aug_27_14_18_25_580517
14:14:00 mriedem then the generation changes twice but we don't log why
14:14:36 zigo mriedem: If you want the full logs and everything, you can find it here: https://review.openstack.org/#/c/597175/
14:14:58 zigo That commit just switches repo from queens to rocky for Debian and puppet-openstack.
14:14:58 mriedem any of those failed jobs?
14:15:02 zigo Yep.
14:15:05 zigo The first one for example.
14:15:33 mriedem yup so in this case we create the RP here http://logs.openstack.org/75/597175/1/check/puppet-openstack-integration-4-scenario001-tempest-debian-stable-luminous/fd38fcf/logs/nova/nova-compute.txt.gz#_2018-08-28_17_09_40_686
14:16:48 mriedem here we POST to placement http://logs.openstack.org/75/597175/1/check/puppet-openstack-integration-4-scenario001-tempest-debian-stable-luminous/fd38fcf/logs/nova/nova-placement-api.txt.gz#_2018-08-28_17_09_40_683
14:16:57 zigo mriedem: What I can do is artificially change the default 0.0 allocation ratio in the package. Would you advise me to do that?
14:17:18 mriedem here we add inventory http://logs.openstack.org/75/597175/1/check/puppet-openstack-integration-4-scenario001-tempest-debian-stable-luminous/fd38fcf/logs/nova/nova-placement-api.txt.gz#_2018-08-28_17_09_40_794
14:18:08 mriedem zigo: i would probably not recommend that at this time no
14:18:19 mriedem i personally would like to figure out what we're dealing with first
14:18:29 cdent agreed
14:18:50 cdent as far as I can tell the nova-compute is never logging when it sets inventory?
14:18:59 mriedem no, it used to
14:20:12 openstackgerrit Merged openstack/osc-placement master: Add image link in README.rst https://review.openstack.org/586839
14:20:25 efried stephenfin: Done (though I'm not a stable core): https://review.openstack.org/597421
14:20:52 stephenfin efried: Ah, indeed. Good enough though
14:21:51 dansmith mriedem: we did discuss having the compute node not override the allocation ratio set on a provider, only use that value when creating it
14:21:55 dansmith did that ever happen?
14:22:07 dansmith I don't see those extra values in conf/compute
14:22:18 mriedem i don't remember that happening no
14:22:38 dansmith https://review.openstack.org/#/c/552105/
14:23:21 mriedem yup, i have unanswered questions in there
14:23:36 dansmith yep, just offering evidence that it didn't
14:24:05 mriedem efried: would be helpful if provider tree's _update_generation method took an "operation" param or something,
14:24:12 mriedem i.e. we're updating the generation b/c inventory was updated or something
14:24:22 sean-k-mooney mriedem: i might add a live migration item to the nova-neutron ptg slot.
14:24:45 sean-k-mooney mriedem: i have stated to address the neutron bugs i have found but looks like there are other nova ones too
14:24:53 efried kosamara: I started going through the spec to make some content, but found it's really going to be easier if we can both post our updates to gerrit.
14:25:08 efried kosamara: I also wanted to know how much you'd looked at the cyborg project, if at all.
14:25:24 mriedem zigo: let me know if/when you have a nova bug posted and i can push some debug patches
14:25:59 mriedem so,
14:26:11 efried mriedem: I would expect the report client side to be logging what it's doing there.
14:26:12 openstack Launchpad bug 1789654 in OpenStack Compute (nova) "placement allocation_ratio initialized with 0.0" [Undecided,New]
14:26:12 zigo mriedem: Yeah, here: https://bugs.launchpad.net/nova/+bug/1789654
14:26:17 mriedem my guess is that update_from_provider_tree is somehow not pushing inventory changes which include the updated allocation_ratio b/c of a cache
14:26:43 openstackgerrit Merged openstack/osc-placement master: Resource provider examples https://review.openstack.org/553461
14:26:50 zigo mriedem: I can add some debian specific patches to my package, and rerun puppet, if you like.
14:27:06 efried mriedem: Ahem, libvirt's update_provider_tree is not setting allocation ratio.
14:27:10 mriedem efried: i know
14:27:12 mriedem the RT does
14:27:15 mriedem see the _normalize method
14:27:16 efried no
14:27:22 zigo RT stands for what?
14:27:28 mriedem ResourceTracker
14:27:29 efried that doesn't get run if you implemented update_provider_tree
14:27:35 mriedem sure it does
14:27:38 efried That said, it should be getting defaulted to 1.0, per handler/inventory
14:27:44 mriedem https://github.com/openstack/nova/blob/6522ea3ecfe99cca3fb33258b11e5a1f34e6e8f0/nova/compute/resource_tracker.py#L901
14:28:15 mriedem and we don't log anything in the report client when we go to flush provider tree inventory changes if the provider tree doesn't think inventory has changed
14:28:17 efried oh wtf...
14:28:57 efried I totally traced this exact thing like last week and was sure we weren't hitting that. /me needs to rework a thing...
14:28:58 mriedem if we didn't call _normalize_inventory_from_cn_obj in this case, libvirt/xen/ironic wouldn't ever have allocation_ratio or reserved values set
14:29:32 openstackgerrit Konstantinos Samaras-Tsakiris proposed openstack/nova-specs master: Placement model for passthrough devices https://review.openstack.org/591037
14:29:32 mriedem and it looks like in some racey cases we never update the inventory b/c the provider tree cache never thinks there is a change
14:29:34 efried edmondsw: Redundant goofy code alert --^
14:30:13 kosamara efried: No problem, we'll coordinate.
14:32:12 efried mriedem: Are you saying it's "normal" for us to set alloc ratio to 0.0 for a sec, because we expect the next update to hit _normalize and then push the updated inventory?
14:32:19 kosamara efried: I had looked at Cyborg in spring, but haven't been following it since. The primary use case that I saw Cyborg enabling was FPGA function-aaS.
14:32:51 mriedem efried: i don't know what is pushing the initial inventory data
14:32:54 mriedem so i can't really say
14:32:54 efried kosamara: Yes, that's kind of the initial motivator, but the project is supposed to subsume all device management.
14:32:56 mriedem we don't log anything
14:33:28 mriedem so i'm going to push some debug patches so we can .... debug
14:33:43 mriedem my guess is the local provider tree cache is current but remote is not, and we never update
14:33:43 efried ack, let me know if I can help.
14:39:50 efried mriedem: Well, looking at has_inventory_changed, it *should* be paying attention to alloc ratio updates, whether the field exists and is being changed, or doesn't exist and is being added.
14:40:10 efried and has_inventory_changed being false is the only way we would avoid sending the update down to placement.
14:40:33 efried ...barring actual errors which we would see in the compute log
14:43:43 openstackgerrit Matt Riedemann proposed openstack/nova master: Log the operation when updating generation in ProviderTree https://review.openstack.org/597553
14:44:10 efried mriedem: FYI: http://paste.openstack.org/show/729053/ <== looks fine.
14:47:08 kosamara efried: Does that mean that instead of proposing improvements to the way Nova manages devices, we should contribute this functionality to Cyborg, because in the future Cyborg will have exclusive responsibility in that domain?
14:47:24 efried kosamara: Well, that's what needs to be discussed, really.

Earlier   Later