Earlier  
Posted Nick Remark
#openstack-nova - 2018-11-01
18:52:24 sean-k-mooney then it would add child resouce providers to the tree created by nova
18:52:42 sean-k-mooney but not modify any resouce provider it did not created
18:57:33 sean-k-mooney efried: today are we doing the provider tree update by put or patch. if put what would it take to make it a patch so nova can do a partial update and merge it on the placement side
18:58:27 efried mriedem: quick fix pls
18:59:29 efried sean-k-mooney: patch is only applicable if you're talking about modifying part of a single provider. Which I think we're not considering.
18:59:56 efried sean-k-mooney: IIUC, you're suggesting modifying some providers in the tree, but not others. That's still PUT - one per provider to be modified.
19:00:05 efried And it's what we do today, see update_from_provider_tree
19:00:17 sean-k-mooney efried: ok cool
19:00:27 openstackgerrit Matt Riedemann proposed openstack/nova master: Remove TODO from get_provider_tree_and_ensure_root https://review.openstack.org/614835
19:00:51 efried +2 ^
19:03:31 sean-k-mooney what i was actully suggesting was make it a single patch call to placement to update all resouce providers in a tree owned by service x but thats a different conversation
19:03:46 efried totally
19:04:48 sean-k-mooney i am still not aware of an usecase that woudl require nova to be aware of a resouce provider created by another service by the way
19:05:10 sean-k-mooney when i say nova i sepcfically mean the compute agent
19:07:59 dansmith efried: so was there some outcome?
19:08:21 efried dansmith: Remember that thing they did where they disabled the refresh_associations?
19:08:51 dansmith oh they reverted that?
19:08:54 dansmith accidentally
19:08:59 efried dansmith: That's why all those calls were zeroes in queens and nonzero once they upgraded (because that hack was no longer there). Yeah.
19:09:04 dansmith ah cool.
19:09:17 efried So belmiro is going to reinstate that and come back at us.
19:09:23 dansmith right on
19:09:32 efried but it's still a shit ton of calls
19:10:06 efried Matt and Sean and I brainstormed briefly on whether we could just get rid of the cache completely (and whether that would actually help).
19:10:22 efried And what we actually ended up doing was getting rid of a comment: https://review.openstack.org/614835 :(
19:13:04 sean-k-mooney efried: well i think we could maybe get rid of the cache but i think it needs a spec not irc ideas
19:13:24 efried sean-k-mooney: I would rather see a PoC in code for that one than a spec.
19:14:11 sean-k-mooney efried: well that a possiblity but it would be similar to neutrons notifications for prot/network events
19:15:14 efried yeah, a subscribable notification framework at placement itself would be cool.
19:15:46 sean-k-mooney yep that is what i was about to type but had not decided how to phase it
19:15:46 mriedem there was a blueprint for that at one point i think
19:16:07 mriedem https://blueprints.launchpad.net/nova/+spec/placement-notifications
19:16:13 sean-k-mooney well this is a much more positive resonce to this idea then i had expected
19:17:29 sean-k-mooney if we had an owner attribute on every provider and a api to register owner with call backs then each service could register a subsciption to the rps they own
19:19:20 mriedem and if ifs and buts were candy and nuts we'd all have a merry christmas
19:19:35 efried I was thinking much simpler to start.
19:19:49 efried You could register for a notification any time $rp_uuid is touched.
19:20:11 efried which includes "create a resource provider with $rp_uuid as a root", which solves the use case we were discussing.
19:20:20 sean-k-mooney efried: i considered that but that could be a lot of RPs
19:20:32 efried Only one per host
19:20:45 sean-k-mooney that said i guess you would only have to do that on creating the rp the first time
19:20:51 sean-k-mooney at least for clean deployments
19:20:52 cfriesen sean-k-mooney: does nova-compute need to know about the child resource providers that cyborg created?
19:21:05 sean-k-mooney cfriesen: i dont think so
19:21:24 sean-k-mooney cfriesen: not in any of the interaction specs i have seen
19:21:25 efried cfriesen: Yeah, we talked about that above; the virt driver needs to know about them for purposes of whitelisting, deploying/attaching, etc.
19:21:39 sean-k-mooney efried: no it does not
19:21:48 sean-k-mooney the whitelisting is cyborges job
19:22:33 sean-k-mooney and deploy/attaching will be not dont in term of the VARs or whatever the equivalent of a port bining has become
19:22:59 cfriesen assuming nova owns a specific set of resource providers, and it's the only thing consuming from those resource providers (can we assume that) then it should only have to update inventory once at startup.
19:23:03 openstackgerrit Merged openstack/os-vif master: Do not import pyroute2 on Windows https://review.openstack.org/614728
19:23:05 ldau Hi, somebody has installed all-in-one openstack using vmware as hypervisor?
19:23:38 sean-k-mooney cfriesen: that is the assumtion we stated in denver so yes i think that is still ture
19:23:39 cfriesen now if anything else can consume those resources, then we need the periodic inventory update
19:23:46 efried sean-k-mooney, cfriesen: if all of that is true, then we can indeed resolve the TODO that mriedem just blew away.
19:24:13 efried cfriesen: You mocking this up?
19:24:24 cfriesen not me. :)
19:24:52 efried cfriesen: "Consume the resources" doesn't matter. Changing inventory matters, but only for the providers I onw.
19:24:53 efried own
19:25:06 efried we don't cache allocation/usage data.
19:25:44 cfriesen efried: agreed. I guess it'd have to be something like CPU/RAM hotplug where it actually changes the inventory
19:26:11 efried But that would be noticed by the virt driver, which would update_provider_tree, which the rt would flush back.
19:26:13 sean-k-mooney cfriesen: it would be the virtdirver that would do that however
19:26:28 efried IOW, we're not getting rid of the cache. We're getting rid of all the cache *refreshing*.
19:26:39 sean-k-mooney efried: yes
19:26:46 efried sean-k-mooney: you mocking this up?
19:26:52 efried or is mriedem?
19:27:19 mriedem i stopped paying attention, what now?
19:27:20 sean-k-mooney efried: if we just disable the refesh in the config dose that not effectivly mock it up
19:27:51 efried mriedem: We're operating on the hypothesis that nova does *not* in fact need to know if outside agents create child providers that they will continue to own.
19:28:06 efried mriedem: And if that's true, we do *not* in fact need to refresh the cache, ever.
19:28:20 efried mriedem: So we *can* in fact resolve the TODO you just removed.
19:28:42 efried sean-k-mooney: More or less, yeah. Which is what CERN did. Which they seem to have had success with.
19:28:50 mriedem how about someone poop this out in the ML and get it sorted out there when gibi, jaypipes and cdent can also weigh in
19:29:00 efried sean-k-mooney: I think there's more we can do, though.
19:29:27 mriedem i think in general, we should default to *not* cache b/c of the cern issue, and only allow caching if you want to opt-in b/c you have a wildly busy env where inventory is changing a lot
19:29:29 efried mriedem: What was your idea to get the compute RP creation happening only once?
19:29:31 mriedem which i guess is a powervm thing
19:29:38 sean-k-mooney efried: yes proably
19:29:59 mriedem efried: yes, create the compute node root rp when the compute node is created, otherwise don't attempt to do the _ensure_resource_provider thing again
19:30:45 mriedem i.e. is it a powervm thing to be swapping out disk and such on the fly and expect nova-compute to just happily handle that?
19:30:57 cfriesen mriedem: changing inventory, or changing usage?
19:31:02 mriedem inventory
19:31:17 mriedem anyway, that probably doesn't matter here,
19:31:27 mriedem we get the inventory regardless to know if it changed so we can push updates back to placement
19:31:27 cfriesen do we expect anyone to have wildly changing inventory? I would have thought inventory is relatively stable.
19:31:35 mriedem cfriesen: that's what i said about 2 hours ago
19:31:37 sean-k-mooney mriedem: if its somthing that is discoverd by the virt driver its not an issue if it chagnes
19:32:26 mriedem i also don't think we really need to worry about refreshing aggregate relationships in the compute,
19:32:29 efried mriedem: I contend it doesn't matter if powervm (or any driver) changes inventory every single periodic. Because update_provider_tree is getting whatever the previous state of placement was (because placement isn't changed yet) and then update_from_provider_tree is flushing that back to placement *and* updating the cache accordingly.
19:32:30 mriedem since we don't do anything with those yet
19:32:38 efried yeah, what sean-k-mooney said, only bigger.
19:33:38 sean-k-mooney so do we all agree we dont need to refesh the cache in any case that is atleas emitly obvious to us
19:33:53 efried tentatively yes
19:34:35 mriedem i would have thought a lot of prior discussion about why we even have a cache in the first place has happene
19:34:37 mriedem *happened
19:34:57 mriedem therefore it seems pretty severe to just all of a sudden say, "oh i guess we don't"
19:35:19 mriedem and if there are reasons, do those reasons justify us caching by default
19:35:35 mriedem anyway, those are questions for the ML, not irc
19:35:37 sean-k-mooney if so then should we start with a patch to default the refresh to off. and a mailing list post to see what operators think/other feedback

Earlier   Later