Earlier  
Posted Nick Remark
#openstack-nova - 2018-11-01
17:33:28 dansmith efried: ^
17:33:39 efried how far back am I reading?
17:33:44 dansmith efried: one line
17:33:52 dansmith and the url he posted
17:35:08 dansmith belmoreira: is it really increasing, or is that you bringing nodes on over time?
17:36:00 belmoreira the increase of requests shows the compute nodes being upgraded over time (queens -> rocky)
17:36:10 dansmith okay
17:36:15 dansmith (ouch)
17:36:22 efried What's happening at that cat-head bump?
17:36:31 efried or possibly batman
17:36:34 dansmith online migrations?
17:37:22 efried those look like trait requests
17:37:34 efried if I'm reading this right.
17:37:57 belmoreira no, it must be a cell that upgraded and then stopped nova-compute
17:38:31 belmoreira so, in the second graph we can see all the new requests
17:39:25 belmoreira UUID/trais ; ?in_tree; UUID/aggregates; ...
17:40:54 efried right, so it looks to me like, before rocky, we weren't calling ?in_tree, UUID/aggregates, or ?member_of at all. Which makes sense.
17:41:27 belmoreira yes, and this seems to be the reason of the increase of requests
17:42:10 efried but also increased number of requests for inventories.
17:42:16 belmoreira but is a huge increase. Just added another graph with the response time of my placement infrastructure
17:42:56 efried I have to say, this isn't all that surprising.
17:44:39 efried although, hm, I would have expected this jump in queens
17:44:57 efried belmoreira: Was this an upgrade from queens, or from earlier?
17:45:33 belmoreira efried from queens
17:47:46 belmoreira I could handle it creating more placement nodes (x3). But looks too much...
17:48:15 efried belmoreira: Can you give me a sense of what this timeline represents? At what point are all the upgrades done and the cloud in stable state?
17:50:34 belmoreira efried the nova/placement control plane was upgraded between 8:00 and 9:00. ~12:00 the compute nodes started to upgrade (this takes 24h for all of them upgrade)
17:51:28 belmoreira at 12:00 (today) almost all compute nodes are in Rocky.
17:51:58 efried belmoreira: So where it tails off at the end, that's when the upgrades are pretty much done?
17:52:02 belmoreira the load graphs shows when I added more capacity for placement
17:52:13 efried Do you have a graph for what it looks like right now?
17:52:30 efried I'm just wondering if it's a massive spike during upgrade, but then it evens back out afterward.
17:52:35 efried in which case... yeah
17:54:45 belmoreira efried I'm getting a new graph from now
17:55:21 efried though once again, I wouldn't have expected e.g. ?in_tree to be zero at queens. That should be happening every periodic.
17:57:42 dansmith efried: you mean you think it's startup storm?
17:57:50 dansmith so every time they reboot computes they'll get this?
17:58:10 efried dansmith: If you reboot a thousand computes...
17:58:26 efried dansmith: I just wanted to understand *whether* it was startup storm.
17:58:27 dansmith right but presumably they're not rebooting them every second
17:58:31 dansmith ack
17:58:39 efried Whether it goes back to normal once everything stabilizes
17:58:41 openstackgerrit Merged openstack/nova stable/rocky: De-dupe subnet IDs when calling neutron /subnets API https://review.openstack.org/608336
17:58:43 dansmith they also know what upgrades look like
17:58:52 efried (I don't)
17:58:56 dansmith so the fact that they're concerned probably means something
17:59:15 efried Heh, I'm not tryng to weasel out of anything. Just trying to grok the problem domain.
17:59:24 dansmith no, I know
17:59:28 dansmith just sain'
18:00:06 dansmith even if we just made the reboot storm a lot worse, that's something we probably need to look at
18:00:12 belmoreira efried a new graph from now
18:00:43 belmoreira it is flat at the end. That is the total number of requests that we handle now
18:01:18 efried dansmith: Can you sanity-check me on this, though - the _refresh_associations code is in queens, including _ensure_resource_provider invoking _get_provider_in_tree, which is what invokes the ?in_tree URI.
18:01:40 efried the mystery being, why would they be seeing zero ?in_tree calls right before the upgrade?
18:01:42 dansmith I just headed into a meeting I have to pay attention to
18:04:34 belmoreira humm. tssurya just point out the "resource_provider_association_refresh" configuration that we had in queens we don't have it in rocky
18:05:15 efried mm, that'd explain a lot. Y'all added that to compensate for this kind of spike in placement traffic iirc
18:05:40 belmoreira efried that explains " I wouldn't have expected e.g. ?in_tree to be zero at queens"
18:05:55 efried yup
18:06:23 mriedem i thought you totally nuked resource_provider_association_refresh rather than just set it to a large value?
18:06:36 efried but also why all those things are zero before the upgrade and nonzero after. Like I was saying, I expect all this stuff to happen at the queens boundary, not rocky.
18:07:58 belmoreira in queens we patch it and set it to a very large number (to not run again). And I miss it now. My fault!
18:08:09 efried IOW I suspect that turning that you would have seen the same graphs simply by turning that switch off and leaving your nodes at queens
18:08:52 belmoreira but the number of requests we really impressive! meaning that is very difficult to keep this option in a large infrastructure
18:09:20 efried belmoreira: I don't disagree with that.
18:10:07 mriedem so by default, every compute (70K?) is refreshing inventory every 1 minute, and every 5 minutes it's also refreshing in_tree, aggregates and traits?
18:10:08 efried I would think moving it to a fairly generous interval and hoping your computes don't all hit that interval at the same time :)
18:10:20 tssurya mriedem: yea
18:10:35 mriedem and we do'nt use the aggregates stuff in compute yet at all from what i can tell
18:10:45 mriedem it was there for sharing providers which we don't support yet
18:10:51 efried well, didn't we start cloning host azs ?
18:10:58 mriedem that's in the API
18:11:18 efried but we're not using that in the scheduler yet?
18:11:43 mriedem the mirrored aggregates stuff? yes there are pre-request placement filters that rely on it (or something external doing the mirroring)
18:11:56 mriedem i'm not sure what that has to do with the cache / refresh for aggregates in all the computes
18:12:39 mriedem iow, i'm not sure what the cache in the compute buys us
18:12:45 efried yeah, I'm actually trying to think what we actually use the cache for at all... right.
18:13:50 efried cdent has been grousing for a while that we should just be able to make placement calls when we need 'em.
18:14:14 sean-k-mooney if we really wanted to make the storm less likely we could use a random prime ofset for the update
18:14:43 efried I was thinking it, but then you said it.
18:15:10 mriedem oslo already does something like that for periodics
18:15:28 efried Trying to think what it would take to rip out the cache completely.
18:15:38 belmoreira I'm changing this option, it will take ~2h to propagate. Will let you know the result
18:15:44 efried ack
18:16:00 belmoreira I have to leave now for some minutes. Thanks for all your help
18:16:22 efried o/
18:17:34 efried mriedem: We use the cache data so the virt driver has the opportunity every periodic to update the provider tree.
18:18:19 efried mriedem: assuming stable placement, no _refresh'ing, we would be doing a helluva lot fewer calls
18:18:31 efried and that's also why we cache agg data. Because upt gets to muck with that stuff also.
18:18:45 mriedem but nothing does right now right?
18:18:48 mriedem for aggregates
18:19:00 mriedem and assuming inventory isn't wildly changing on a compute node, we don't really need the cache
18:19:49 efried not sure I'm following.
18:20:03 efried are you saying "as long as nothing is changing, we don't need to call update_provider_tree" ?
18:20:25 mriedem update_provider_tree is what returns the inventory from the driver to the RT to push off to placement every 60 seocnds
18:20:27 mriedem *seconds
18:20:28 mriedem right?
18:21:04 efried Yes
18:21:05 mriedem and assuming that disk/ram/cpu on a host doesn't change all that often, at least without a restart of the host, it seems odd we need to cache that information
18:21:18 efried But how else would we know whether to push the info back to placement?

Earlier   Later