Earlier  
Posted Nick Remark
#openstack-nova - 2018-11-01
17:42:16 belmoreira but is a huge increase. Just added another graph with the response time of my placement infrastructure
17:42:56 efried I have to say, this isn't all that surprising.
17:44:39 efried although, hm, I would have expected this jump in queens
17:44:57 efried belmoreira: Was this an upgrade from queens, or from earlier?
17:45:33 belmoreira efried from queens
17:47:46 belmoreira I could handle it creating more placement nodes (x3). But looks too much...
17:48:15 efried belmoreira: Can you give me a sense of what this timeline represents? At what point are all the upgrades done and the cloud in stable state?
17:50:34 belmoreira efried the nova/placement control plane was upgraded between 8:00 and 9:00. ~12:00 the compute nodes started to upgrade (this takes 24h for all of them upgrade)
17:51:28 belmoreira at 12:00 (today) almost all compute nodes are in Rocky.
17:51:58 efried belmoreira: So where it tails off at the end, that's when the upgrades are pretty much done?
17:52:02 belmoreira the load graphs shows when I added more capacity for placement
17:52:13 efried Do you have a graph for what it looks like right now?
17:52:30 efried I'm just wondering if it's a massive spike during upgrade, but then it evens back out afterward.
17:52:35 efried in which case... yeah
17:54:45 belmoreira efried I'm getting a new graph from now
17:55:21 efried though once again, I wouldn't have expected e.g. ?in_tree to be zero at queens. That should be happening every periodic.
17:57:42 dansmith efried: you mean you think it's startup storm?
17:57:50 dansmith so every time they reboot computes they'll get this?
17:58:10 efried dansmith: If you reboot a thousand computes...
17:58:26 efried dansmith: I just wanted to understand *whether* it was startup storm.
17:58:27 dansmith right but presumably they're not rebooting them every second
17:58:31 dansmith ack
17:58:39 efried Whether it goes back to normal once everything stabilizes
17:58:41 openstackgerrit Merged openstack/nova stable/rocky: De-dupe subnet IDs when calling neutron /subnets API https://review.openstack.org/608336
17:58:43 dansmith they also know what upgrades look like
17:58:52 efried (I don't)
17:58:56 dansmith so the fact that they're concerned probably means something
17:59:15 efried Heh, I'm not tryng to weasel out of anything. Just trying to grok the problem domain.
17:59:24 dansmith no, I know
17:59:28 dansmith just sain'
18:00:06 dansmith even if we just made the reboot storm a lot worse, that's something we probably need to look at
18:00:12 belmoreira efried a new graph from now
18:00:43 belmoreira it is flat at the end. That is the total number of requests that we handle now
18:01:18 efried dansmith: Can you sanity-check me on this, though - the _refresh_associations code is in queens, including _ensure_resource_provider invoking _get_provider_in_tree, which is what invokes the ?in_tree URI.
18:01:40 efried the mystery being, why would they be seeing zero ?in_tree calls right before the upgrade?
18:01:42 dansmith I just headed into a meeting I have to pay attention to
18:04:34 belmoreira humm. tssurya just point out the "resource_provider_association_refresh" configuration that we had in queens we don't have it in rocky
18:05:15 efried mm, that'd explain a lot. Y'all added that to compensate for this kind of spike in placement traffic iirc
18:05:40 belmoreira efried that explains " I wouldn't have expected e.g. ?in_tree to be zero at queens"
18:05:55 efried yup
18:06:23 mriedem i thought you totally nuked resource_provider_association_refresh rather than just set it to a large value?
18:06:36 efried but also why all those things are zero before the upgrade and nonzero after. Like I was saying, I expect all this stuff to happen at the queens boundary, not rocky.
18:07:58 belmoreira in queens we patch it and set it to a very large number (to not run again). And I miss it now. My fault!
18:08:09 efried IOW I suspect that turning that you would have seen the same graphs simply by turning that switch off and leaving your nodes at queens
18:08:52 belmoreira but the number of requests we really impressive! meaning that is very difficult to keep this option in a large infrastructure
18:09:20 efried belmoreira: I don't disagree with that.
18:10:07 mriedem so by default, every compute (70K?) is refreshing inventory every 1 minute, and every 5 minutes it's also refreshing in_tree, aggregates and traits?
18:10:08 efried I would think moving it to a fairly generous interval and hoping your computes don't all hit that interval at the same time :)
18:10:20 tssurya mriedem: yea
18:10:35 mriedem and we do'nt use the aggregates stuff in compute yet at all from what i can tell
18:10:45 mriedem it was there for sharing providers which we don't support yet
18:10:51 efried well, didn't we start cloning host azs ?
18:10:58 mriedem that's in the API
18:11:18 efried but we're not using that in the scheduler yet?
18:11:43 mriedem the mirrored aggregates stuff? yes there are pre-request placement filters that rely on it (or something external doing the mirroring)
18:11:56 mriedem i'm not sure what that has to do with the cache / refresh for aggregates in all the computes
18:12:39 mriedem iow, i'm not sure what the cache in the compute buys us
18:12:45 efried yeah, I'm actually trying to think what we actually use the cache for at all... right.
18:13:50 efried cdent has been grousing for a while that we should just be able to make placement calls when we need 'em.
18:14:14 sean-k-mooney if we really wanted to make the storm less likely we could use a random prime ofset for the update
18:14:43 efried I was thinking it, but then you said it.
18:15:10 mriedem oslo already does something like that for periodics
18:15:28 efried Trying to think what it would take to rip out the cache completely.
18:15:38 belmoreira I'm changing this option, it will take ~2h to propagate. Will let you know the result
18:15:44 efried ack
18:16:00 belmoreira I have to leave now for some minutes. Thanks for all your help
18:16:22 efried o/
18:17:34 efried mriedem: We use the cache data so the virt driver has the opportunity every periodic to update the provider tree.
18:18:19 efried mriedem: assuming stable placement, no _refresh'ing, we would be doing a helluva lot fewer calls
18:18:31 efried and that's also why we cache agg data. Because upt gets to muck with that stuff also.
18:18:45 mriedem but nothing does right now right?
18:18:48 mriedem for aggregates
18:19:00 mriedem and assuming inventory isn't wildly changing on a compute node, we don't really need the cache
18:19:49 efried not sure I'm following.
18:20:03 efried are you saying "as long as nothing is changing, we don't need to call update_provider_tree" ?
18:20:25 mriedem update_provider_tree is what returns the inventory from the driver to the RT to push off to placement every 60 seocnds
18:20:27 mriedem *seconds
18:20:28 mriedem right?
18:21:04 efried Yes
18:21:05 mriedem and assuming that disk/ram/cpu on a host doesn't change all that often, at least without a restart of the host, it seems odd we need to cache that information
18:21:18 efried But how else would we know whether to push the info back to placement?
18:21:51 mriedem in the before upt times, didn't the RT/report client just pull inventory, compare to what was reported by the driver, and the PUT it back if there were changes?
18:22:08 efried What does "pull inventory" mean, though?
18:22:22 efried pull from placement
18:22:28 mriedem GET /resource_providers/{rp_uuid}/inventories
18:22:29 efried i.e. GET /rps/UUID/inventory
18:22:32 efried yeah
18:22:42 sean-k-mooney efried: well the driver could have a perodic check but rememebr the last value it sent and only send a value if it detactes there was a chage
18:22:50 efried sean-k-mooney: ^ cache
18:22:58 sean-k-mooney that not the same as a cache
18:23:00 efried and that's what we do
18:23:39 mriedem get_provider_tree_and_ensure_root is what gets the provider tree from the report client and pulls the current inventory from placement, yes?
18:23:48 efried yes
18:23:48 mriedem and also checks to see that the provider exists on every periodic
18:23:51 mriedem which we should actually know
18:24:37 efried yeah, we could conceivably expect the compute RP not to disappear once we've created it.
18:24:45 efried I mean, I don't know how resilient we're trying to be in the face of OOB changes.
18:24:59 efried we do offer a placement CLI, not just for GETs but for writes as well
18:25:47 mriedem the compute service record, compute node record, and rp can all be deleted if the compute service record is deleted
18:25:56 mriedem but to get the compute service record back, you have to restart the compute service

Earlier   Later