| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2019-12-03 | |||
| 18:13:49 | dansmith | sweet, so that's likely placement yeah | |
| 18:13:54 | eandersson | Also scaling computes | |
| 18:14:05 | dansmith | eandersson: what does "scaling computes" mean? | |
| 18:14:06 | eandersson | We initially limited ourself to 1000 computes per region | |
| 18:14:12 | eandersson | in Mitaka | |
| 18:14:23 | eandersson | but now believe we can reach a much higher number than that | |
| 18:14:55 | dansmith | eandersson: is that via cells or just in general? because I don't think we've done anything other than get *more* chatty to the db/mq since mitaka :) | |
| 18:14:57 | eandersson | And even if we do hit limitations we can now use cells | |
| 18:15:23 | eandersson | scheduler used to be really heavy on rmq | |
| 18:15:29 | eandersson | in mitaka | |
| 18:15:53 | sean-k-mooney | eandersson: do you run sepperate rmq instacne per openstack service | |
| 18:16:01 | eandersson | We do not | |
| 18:16:24 | eandersson | but even at 1k computes we are barely putting any stress on rmq at the moment | |
| 18:16:32 | eandersson | Of course we are not using ceilometer | |
| 18:16:34 | sean-k-mooney | im surpirsed you are getting to 1000 nodes on one cluster with neutron and nova sharing it | |
| 18:16:48 | dansmith | eandersson: so you're thinking that lowered load on rabbit from the scheduler lets you have more computes? | |
| 18:16:52 | sean-k-mooney | ya not using ceilometer helps | |
| 18:17:02 | eandersson | dansmith I think it helps | |
| 18:17:09 | dansmith | ack | |
| 18:18:12 | eandersson | We were getting a lot of slow api calls in mitaka and they are all very consistent now | |
| 18:19:06 | eandersson | The only problems we have with nova now is getting our super custom scheduling logic to scale | |
| 18:19:07 | dansmith | that's what I'm interested in specifically, | |
| 18:19:17 | dansmith | but without knowing which calls those were I can't really attribute them to anything | |
| 18:19:31 | dansmith | (that == slow api calls) | |
| 18:20:55 | eandersson | Yea - unfortunately a lot of the research and testing we did was back in ~2016 | |
| 18:21:48 | eandersson | We didn't do a great job tracking individual improvements when going from Mitaka to Rocky | |
| 18:21:53 | sean-k-mooney | eandersson: do you have a list of constraits you need to schduler on that are not supported cleanly upstream that could be shared perhaps we could accomadate some of your custom logic | |
| 18:22:44 | eandersson | We do flavor stacking and what... we internally call "perfect fit". | |
| 18:22:59 | sean-k-mooney | using the type affingity filter | |
| 18:23:07 | mriedem | you also have a variant of the old flavor affinity filter yeah? | |
| 18:23:08 | sean-k-mooney | so that each host only one flavor | |
| 18:23:30 | eandersson | So we actually want to stack flavors on a compute | |
| 18:23:34 | eandersson | because these are game servers | |
| 18:23:40 | eandersson | So each game server takes up one numa | |
| 18:24:07 | eandersson | but we still want to be able to schedule other micro services on top | |
| 18:24:50 | eandersson | Since we don't want to have to divide the fleet | |
| 18:24:55 | sean-k-mooney | so you want to pack the large flavor and then fit the micof servces wehre tehy can | |
| 18:24:59 | eandersson | Yep | |
| 18:25:17 | sean-k-mooney | that not really unresobaly to be fair | |
| 18:25:41 | sean-k-mooney | the main issue i guess you face right now is fragmenation | |
| 18:26:02 | eandersson | Yea - if we go with the out of the box implementation | |
| 18:26:04 | sean-k-mooney | e.g. a small instnace spawns preventing a large instnace | |
| 18:26:09 | eandersson | yep | |
| 18:26:13 | dansmith | but also...scheduler filters are the one place I think we *should* be pluggable, and so unless other people want *exactly* the same weird scheduling thing, it makes sense for them to do this on their own, IMHO | |
| 18:26:37 | eandersson | Yep - I hate that I have to patch nova for this | |
| 18:26:55 | eandersson | I mean we have a super custom way of deploying, plus probably 20 custom nova patches at least | |
| 18:27:02 | sean-k-mooney | this is not the first time i have heard this requrest however | |
| 18:27:17 | eandersson | So it's not a big deal, but would be a lot easier for us to manage it if it was pluggable | |
| 18:27:27 | dansmith | scheduler filters *are* pluggable | |
| 18:27:35 | sean-k-mooney | eandersson: as are the weighers | |
| 18:27:36 | dansmith | presumably you're dependent on other changes/ | |
| 18:27:46 | eandersson | They are? | |
| 18:27:50 | sean-k-mooney | yep | |
| 18:28:00 | sean-k-mooney | but its non ovious how to do it | |
| 18:28:09 | dansmith | it's very obvious | |
| 18:28:20 | dansmith | it may not be _documented_ :) | |
| 18:28:34 | eandersson | btw we also have weights for upgrading computes (e.g. a compute that needs a OS upgrade would be moved into an aggregate to reduce the changes of it getting scheduled to) | |
| 18:28:48 | dansmith | https://docs.openstack.org/nova/latest/user/filter-scheduler.html#writing-your-own-filter | |
| 18:29:00 | sean-k-mooney | eandersson: here is an example https://opendev.org/x/nfv-filters | |
| 18:30:15 | eandersson | Ah yea I see | |
| 18:31:08 | sean-k-mooney | you just do filter_scheduler.available_filters=nova.scheduler.filters.all_filters,nfv_filters.nova.scheduler.filters.aggregate_instance_type_filter | |
| 18:31:15 | openstackgerrit | Stephen Finucane proposed openstack/nova master: nova-net: Kill it https://review.opendev.org/696518 | |
| 18:31:15 | openstackgerrit | Stephen Finucane proposed openstack/nova master: Rename 'nova.network.neutronv2' -> 'nova.network' https://review.opendev.org/696745 | |
| 18:31:16 | openstackgerrit | Stephen Finucane proposed openstack/nova master: Rename 'nova.network.security_group.neutron_driver' -> 'nova.network.security_group' https://review.opendev.org/696746 | |
| 18:31:17 | openstackgerrit | Stephen Finucane proposed openstack/nova master: Remove unnecessary 'neutronv2' prefixes https://review.opendev.org/696776 | |
| 18:31:17 | openstackgerrit | Stephen Finucane proposed openstack/nova master: nova-net: Remove unused exceptions https://review.opendev.org/697149 | |
| 18:31:17 | openstackgerrit | Stephen Finucane proposed openstack/nova master: nova-net: Remove db methods for ProviderMethod https://review.opendev.org/697150 | |
| 18:31:18 | openstackgerrit | Stephen Finucane proposed openstack/nova master: nova-net: Remove unused 'stub_out_db_network_api' https://review.opendev.org/697151 | |
| 18:31:18 | openstackgerrit | Stephen Finucane proposed openstack/nova master: nova-net: Remove remaining nova-network quotas https://review.opendev.org/697152 | |
| 18:31:19 | openstackgerrit | Stephen Finucane proposed openstack/nova master: nova-net: Remove use of legacy 'FloatingIP' object https://review.opendev.org/697153 | |
| 18:31:19 | openstackgerrit | Stephen Finucane proposed openstack/nova master: nova-net: Remove use of legacy 'Network' object https://review.opendev.org/697154 | |
| 18:31:20 | openstackgerrit | Stephen Finucane proposed openstack/nova master: nova-net: Remove use of legacy 'SecurityGroup' object https://review.opendev.org/697155 | |
| 18:31:20 | openstackgerrit | Stephen Finucane proposed openstack/nova master: nova-net: Remove unused nova-network objects https://review.opendev.org/697156 | |
| 18:32:07 | sean-k-mooney | you can similary have out of tree weighers that you can load in a similar way | |
| 18:34:19 | sean-k-mooney | dansmith: by the way this is actully a lot simpler then i rememeberd i was thinking of out of tree virt drivers which require you to reopen the nova namespace and do other things to make them work | |
| 18:34:39 | sean-k-mooney | dansmith: be we dont really want that to be plugablle in the same way so it makes sense that is more work | |
| 18:34:54 | dansmith | yup | |
| 18:49:33 | artom | Do we not reset the old flavor back on the instance if we fail a resize? Fail as in, something goes wrong during _prep_resize or resize_instance | |
| 18:49:51 | efried | gosh, I would hope we do | |
| 18:50:19 | artom | I can't find it - maybe I'm looking in the wrong place... | |
| 18:50:24 | artom | I mean, we must | |
| 18:50:38 | sean-k-mooney | in pre_resize i dont think we have saved the instace yet | |
| 18:51:22 | artom | Oh, is that how? | |
| 18:51:29 | sean-k-mooney | i have not looked at that in a while however. have we saved it since old_flavor was defined | |
| 18:51:30 | artom | We just don't persist until it's final | |
| 18:51:51 | sean-k-mooney | well we woudl presisti it before we go to resize verify | |
| 18:52:44 | sean-k-mooney | i would just check where it is saved first | |
| 18:53:02 | artom | That's what I'm looking for... | |
| 18:53:03 | dansmith | not just save, but save() after instance.flavor is set, unrelated to instance.old_flavor | |
| 18:55:32 | sean-k-mooney | i think this is where we revert it if we get to revert resize https://github.com/openstack/nova/blob/757fc03b78d542e7262343b65eacea02ce11dd04/nova/objects/instance.py#L1021-L1035 | |
| 18:55:44 | sean-k-mooney | but im not sure that is required if we fail early | |
| 18:58:21 | sean-k-mooney | this is where we update the instace i think https://github.com/openstack/nova/blob/757fc03b78d542e7262343b65eacea02ce11dd04/nova/compute/manager.py#L5260-L5292 | |
| 18:59:25 | sean-k-mooney | so assuming we call _prep_resize before _finish_resize if you fail in _prep_resize im not sure you need to do anything | |
| 19:00:28 | sean-k-mooney | artom: is that what you were looking for? | |
| 19:02:54 | artom | sean-k-mooney, I'm looking for where the request spec is reverted back to the old flavor | |
| 19:03:01 | artom | I said instance dind't I? | |
| 19:03:05 | artom | I meant request spec | |
| 19:03:49 | openstackgerrit | Matt Riedemann proposed openstack/nova master: WIP: Implement cleanup_instance_network_on_host for neutron API https://review.opendev.org/697162 | |
| 19:07:12 | dansmith | artom: compute can't reset the reqspec, so if it happens late enough, that won't happen | |
| 19:07:53 | dansmith | er, s/happens/fails/ | |