| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-11-01 | |||
| 18:59:56 | efried | sean-k-mooney: IIUC, you're suggesting modifying some providers in the tree, but not others. That's still PUT - one per provider to be modified. | |
| 19:00:05 | efried | And it's what we do today, see update_from_provider_tree | |
| 19:00:17 | sean-k-mooney | efried: ok cool | |
| 19:00:27 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Remove TODO from get_provider_tree_and_ensure_root https://review.openstack.org/614835 | |
| 19:00:51 | efried | +2 ^ | |
| 19:03:31 | sean-k-mooney | what i was actully suggesting was make it a single patch call to placement to update all resouce providers in a tree owned by service x but thats a different conversation | |
| 19:03:46 | efried | totally | |
| 19:04:48 | sean-k-mooney | i am still not aware of an usecase that woudl require nova to be aware of a resouce provider created by another service by the way | |
| 19:05:10 | sean-k-mooney | when i say nova i sepcfically mean the compute agent | |
| 19:07:59 | dansmith | efried: so was there some outcome? | |
| 19:08:21 | efried | dansmith: Remember that thing they did where they disabled the refresh_associations? | |
| 19:08:51 | dansmith | oh they reverted that? | |
| 19:08:54 | dansmith | accidentally | |
| 19:08:59 | efried | dansmith: That's why all those calls were zeroes in queens and nonzero once they upgraded (because that hack was no longer there). Yeah. | |
| 19:09:04 | dansmith | ah cool. | |
| 19:09:17 | efried | So belmiro is going to reinstate that and come back at us. | |
| 19:09:23 | dansmith | right on | |
| 19:09:32 | efried | but it's still a shit ton of calls | |
| 19:10:06 | efried | Matt and Sean and I brainstormed briefly on whether we could just get rid of the cache completely (and whether that would actually help). | |
| 19:10:22 | efried | And what we actually ended up doing was getting rid of a comment: https://review.openstack.org/614835 :( | |
| 19:13:04 | sean-k-mooney | efried: well i think we could maybe get rid of the cache but i think it needs a spec not irc ideas | |
| 19:13:24 | efried | sean-k-mooney: I would rather see a PoC in code for that one than a spec. | |
| 19:14:11 | sean-k-mooney | efried: well that a possiblity but it would be similar to neutrons notifications for prot/network events | |
| 19:15:14 | efried | yeah, a subscribable notification framework at placement itself would be cool. | |
| 19:15:46 | sean-k-mooney | yep that is what i was about to type but had not decided how to phase it | |
| 19:15:46 | mriedem | there was a blueprint for that at one point i think | |
| 19:16:07 | mriedem | https://blueprints.launchpad.net/nova/+spec/placement-notifications | |
| 19:16:13 | sean-k-mooney | well this is a much more positive resonce to this idea then i had expected | |
| 19:17:29 | sean-k-mooney | if we had an owner attribute on every provider and a api to register owner with call backs then each service could register a subsciption to the rps they own | |
| 19:19:20 | mriedem | and if ifs and buts were candy and nuts we'd all have a merry christmas | |
| 19:19:35 | efried | I was thinking much simpler to start. | |
| 19:19:49 | efried | You could register for a notification any time $rp_uuid is touched. | |
| 19:20:11 | efried | which includes "create a resource provider with $rp_uuid as a root", which solves the use case we were discussing. | |
| 19:20:20 | sean-k-mooney | efried: i considered that but that could be a lot of RPs | |
| 19:20:32 | efried | Only one per host | |
| 19:20:45 | sean-k-mooney | that said i guess you would only have to do that on creating the rp the first time | |
| 19:20:51 | sean-k-mooney | at least for clean deployments | |
| 19:20:52 | cfriesen | sean-k-mooney: does nova-compute need to know about the child resource providers that cyborg created? | |
| 19:21:05 | sean-k-mooney | cfriesen: i dont think so | |
| 19:21:24 | sean-k-mooney | cfriesen: not in any of the interaction specs i have seen | |
| 19:21:25 | efried | cfriesen: Yeah, we talked about that above; the virt driver needs to know about them for purposes of whitelisting, deploying/attaching, etc. | |
| 19:21:39 | sean-k-mooney | efried: no it does not | |
| 19:21:48 | sean-k-mooney | the whitelisting is cyborges job | |
| 19:22:33 | sean-k-mooney | and deploy/attaching will be not dont in term of the VARs or whatever the equivalent of a port bining has become | |
| 19:22:59 | cfriesen | assuming nova owns a specific set of resource providers, and it's the only thing consuming from those resource providers (can we assume that) then it should only have to update inventory once at startup. | |
| 19:23:03 | openstackgerrit | Merged openstack/os-vif master: Do not import pyroute2 on Windows https://review.openstack.org/614728 | |
| 19:23:05 | ldau | Hi, somebody has installed all-in-one openstack using vmware as hypervisor? | |
| 19:23:38 | sean-k-mooney | cfriesen: that is the assumtion we stated in denver so yes i think that is still ture | |
| 19:23:39 | cfriesen | now if anything else can consume those resources, then we need the periodic inventory update | |
| 19:23:46 | efried | sean-k-mooney, cfriesen: if all of that is true, then we can indeed resolve the TODO that mriedem just blew away. | |
| 19:24:13 | efried | cfriesen: You mocking this up? | |
| 19:24:24 | cfriesen | not me. :) | |
| 19:24:52 | efried | cfriesen: "Consume the resources" doesn't matter. Changing inventory matters, but only for the providers I onw. | |
| 19:24:53 | efried | own | |
| 19:25:06 | efried | we don't cache allocation/usage data. | |
| 19:25:44 | cfriesen | efried: agreed. I guess it'd have to be something like CPU/RAM hotplug where it actually changes the inventory | |
| 19:26:11 | efried | But that would be noticed by the virt driver, which would update_provider_tree, which the rt would flush back. | |
| 19:26:13 | sean-k-mooney | cfriesen: it would be the virtdirver that would do that however | |
| 19:26:28 | efried | IOW, we're not getting rid of the cache. We're getting rid of all the cache *refreshing*. | |
| 19:26:39 | sean-k-mooney | efried: yes | |
| 19:26:46 | efried | sean-k-mooney: you mocking this up? | |
| 19:26:52 | efried | or is mriedem? | |
| 19:27:19 | mriedem | i stopped paying attention, what now? | |
| 19:27:20 | sean-k-mooney | efried: if we just disable the refesh in the config dose that not effectivly mock it up | |
| 19:27:51 | efried | mriedem: We're operating on the hypothesis that nova does *not* in fact need to know if outside agents create child providers that they will continue to own. | |
| 19:28:06 | efried | mriedem: And if that's true, we do *not* in fact need to refresh the cache, ever. | |
| 19:28:20 | efried | mriedem: So we *can* in fact resolve the TODO you just removed. | |
| 19:28:42 | efried | sean-k-mooney: More or less, yeah. Which is what CERN did. Which they seem to have had success with. | |
| 19:28:50 | mriedem | how about someone poop this out in the ML and get it sorted out there when gibi, jaypipes and cdent can also weigh in | |
| 19:29:00 | efried | sean-k-mooney: I think there's more we can do, though. | |
| 19:29:27 | mriedem | i think in general, we should default to *not* cache b/c of the cern issue, and only allow caching if you want to opt-in b/c you have a wildly busy env where inventory is changing a lot | |
| 19:29:29 | efried | mriedem: What was your idea to get the compute RP creation happening only once? | |
| 19:29:31 | mriedem | which i guess is a powervm thing | |
| 19:29:38 | sean-k-mooney | efried: yes proably | |
| 19:29:59 | mriedem | efried: yes, create the compute node root rp when the compute node is created, otherwise don't attempt to do the _ensure_resource_provider thing again | |
| 19:30:45 | mriedem | i.e. is it a powervm thing to be swapping out disk and such on the fly and expect nova-compute to just happily handle that? | |
| 19:30:57 | cfriesen | mriedem: changing inventory, or changing usage? | |
| 19:31:02 | mriedem | inventory | |
| 19:31:17 | mriedem | anyway, that probably doesn't matter here, | |
| 19:31:27 | mriedem | we get the inventory regardless to know if it changed so we can push updates back to placement | |
| 19:31:27 | cfriesen | do we expect anyone to have wildly changing inventory? I would have thought inventory is relatively stable. | |
| 19:31:35 | mriedem | cfriesen: that's what i said about 2 hours ago | |
| 19:31:37 | sean-k-mooney | mriedem: if its somthing that is discoverd by the virt driver its not an issue if it chagnes | |
| 19:32:26 | mriedem | i also don't think we really need to worry about refreshing aggregate relationships in the compute, | |
| 19:32:29 | efried | mriedem: I contend it doesn't matter if powervm (or any driver) changes inventory every single periodic. Because update_provider_tree is getting whatever the previous state of placement was (because placement isn't changed yet) and then update_from_provider_tree is flushing that back to placement *and* updating the cache accordingly. | |
| 19:32:30 | mriedem | since we don't do anything with those yet | |
| 19:32:38 | efried | yeah, what sean-k-mooney said, only bigger. | |
| 19:33:38 | sean-k-mooney | so do we all agree we dont need to refesh the cache in any case that is atleas emitly obvious to us | |
| 19:33:53 | efried | tentatively yes | |
| 19:34:35 | mriedem | i would have thought a lot of prior discussion about why we even have a cache in the first place has happene | |
| 19:34:37 | mriedem | *happened | |
| 19:34:57 | mriedem | therefore it seems pretty severe to just all of a sudden say, "oh i guess we don't" | |
| 19:35:19 | mriedem | and if there are reasons, do those reasons justify us caching by default | |
| 19:35:35 | mriedem | anyway, those are questions for the ML, not irc | |
| 19:35:37 | sean-k-mooney | if so then should we start with a patch to default the refresh to off. and a mailing list post to see what operators think/other feedback | |
| 19:35:55 | mriedem | cern is the only deployment big enough and new enough that i've heard complain about that refresh | |
| 19:36:01 | mriedem | not sure if mnaser is doing anything about it | |
| 19:36:20 | mnaser | hi | |
| 19:36:42 | mriedem | mnaser: tl;dr do you turn this way down? https://docs.openstack.org/nova/latest/configuration/config.html#compute.resource_provider_association_refresh | |
| 19:37:01 | mriedem | to avoid computes DDoS'ing placement every 5 minutes | |