| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-05-03 | |||
| 14:09:55 | dansmith | bauzas: the second aspect is that you add a compute, you wait for it to register, you wait for it to be discovered, then you go and add it to the right az | |
| 14:10:16 | dansmith | that's stuff that a deployment tool has to do a bunch of polling and looping over, but it really shouldn't | |
| 14:10:40 | sean-k-mooney | dansmith: if we remvoed the exists check in the host aggrate add | |
| 14:10:45 | dansmith | if it could say "this should be in AZ foo" then all that stuff can happen automatically and without a bunch of slow steps in the process | |
| 14:10:47 | bauzas | dansmith: technically you don't add a host to an az, but to an aggregate | |
| 14:10:55 | sean-k-mooney | and we leveraged the fact we can right a uuid for the comptue to use | |
| 14:11:02 | dansmith | sean-k-mooney: yeah but then people can typo things when they add them, so I don't love that | |
| 14:11:16 | dansmith | bauzas: I know :) | |
| 14:11:20 | bauzas | dansmith: and an operator can add a host to an aggregate as soon as it's registered in the services table | |
| 14:11:20 | sean-k-mooney | i could see the workflow being the installer generate a uuid for the host, then adds it to the righ host aggrate and then deploys the host | |
| 14:11:41 | sean-k-mooney | dansmith: ya that one of the reasons we dont allow it today | |
| 14:11:44 | dansmith | sean-k-mooney: but you don't add compute nodes to aggregates, you add services right? | |
| 14:11:56 | sean-k-mooney | ah you are right its the service | |
| 14:11:56 | bauzas | dansmith: typos can't append, we query the record based on the service name | |
| 14:12:04 | sean-k-mooney | which is why this does nto really work with ironic today | |
| 14:12:05 | bauzas | happen* | |
| 14:12:21 | dansmith | bauzas: sean-k-mooney is suggesting we drop that requirement | |
| 14:12:33 | sean-k-mooney | dansmith: well sahid was previously | |
| 14:12:35 | bauzas | again, why ? | |
| 14:12:53 | sean-k-mooney | bauzas: to allwo you to pregreister the service to host aggrate | |
| 14:12:56 | sean-k-mooney | before its created | |
| 14:13:00 | bauzas | can't a script loop over a service list and do the aggregate add after ? | |
| 14:13:16 | dansmith | I'm just thinking of how customers are sometimes afraid to run stack update on a running cloud, just like I'm afraid to re-run my own shoddy ansible sometimes :) | |
| 14:13:37 | dansmith | and how making the act of adding a new compute more of a compute-focused thing than a deployment-focused thing would likely be a welcome pattern for people | |
| 14:13:50 | bauzas | sean-k-mooney: we would soften the requirements for the very little benefit of pre-adding a compute to an aggregate | |
| 14:14:02 | sean-k-mooney | bauzas: you can but if you do that in cron and a vm is booted that request nova in that interval then it will be broken | |
| 14:14:06 | dansmith | bauzas: I agree, I don't like that idea | |
| 14:14:18 | sean-k-mooney | once an instance gets schdeuled to the host it cant move to a real az | |
| 14:14:46 | bauzas | I like the idea to only add something schedulable if that thing can actually really schedule my stuff | |
| 14:15:01 | sean-k-mooney | could we hide the nova az if all compute service in it were disabled | |
| 14:15:05 | bauzas | the second it got added to an aggregate, it can be scheduled | |
| 14:15:13 | bauzas | again, why ? | |
| 14:15:17 | sean-k-mooney | then you can set the config option to only register them as disabled | |
| 14:15:31 | bauzas | we would put a lot of preconditions and added complexity for a very little gain | |
| 14:15:48 | bauzas | and adding a compute isn't really a daily operation | |
| 14:15:57 | sean-k-mooney | bauzas: sure its little gain but its a ligitmate pain point for there end customer apperenlty | |
| 14:16:07 | sean-k-mooney | well it depend on your scale | |
| 14:16:34 | bauzas | sean-k-mooney: the bug report wasn't complaining about it | |
| 14:16:44 | dansmith | we already have a task that has to run for a new compute (discover) which can either be automated in scheduler, called via cron, or manually during a deployment.. its purpose is to do things in the api database, so it feels like a very natural thing to have that step also do this, which is placing the compute in the right grouping | |
| 14:16:51 | bauzas | sean-k-mooney: the bug report was about to prevent 'nova' to show off | |
| 14:16:58 | sean-k-mooney | bauzas: sahid said it caused confustion for ther custoemr. | |
| 14:17:26 | bauzas | dansmith: I'm confused, everything is API-driven | |
| 14:17:35 | dansmith | sahid: is the confusion of showing the wrong AZ the only thing you care about? | |
| 14:17:47 | dansmith | bauzas: adding a new compute is not api-driven | |
| 14:17:55 | dansmith | bauzas: the exception to that is needing to put it into the right az | |
| 14:18:43 | sean-k-mooney | bauzas: expanding on ^ you can delete compute services from the api, you can only create them via rpc via the conductor | |
| 14:19:48 | sean-k-mooney | bauzas: we do not have a create endpoint https://docs.openstack.org/api-ref/compute/#compute-services-os-services | |
| 14:20:08 | bauzas | OK I admit the feature gap | |
| 14:20:16 | sean-k-mooney | dansmith: addign a create endpoint woudl also work i guess. when the compute starts it would find the existing record | |
| 14:20:40 | sean-k-mooney | and ocne the record is created you could perhaps add the servicew to an aggreate | |
| 14:20:41 | dansmith | sean-k-mooney: to do that we'd have to expose the concept of a cell to the API which I'm -3 on :) | |
| 14:20:41 | bauzas | but I'm very reluctant to make add_host_to_agg() less strict | |
| 14:20:59 | sean-k-mooney | oh ya we would | |
| 14:21:00 | dansmith | making this more manual is also not something I'm in favor of | |
| 14:21:05 | sean-k-mooney | case this is in the cell db | |
| 14:21:12 | sean-k-mooney | so ya thats not an option relaly | |
| 14:21:21 | dansmith | and pre-creating hosts makes that more manual, and much more likely to be wrong, given what we know about hostnames :) | |
| 14:21:34 | sean-k-mooney | very true | |
| 14:21:35 | dansmith | I'd like to make this more automatic | |
| 14:21:49 | dansmith | which may or may not be what sahid really wants, but making it more automatic will *also* solve sahid's concern | |
| 14:21:50 | bauzas | without any upcalls for sure | |
| 14:21:55 | dansmith | yep | |
| 14:22:12 | bauzas | which is quite a chicken-and-egg problem | |
| 14:22:14 | sean-k-mooney | the thing is we would like this to be more atomatic while also not havign upcalls or having the comptue specificly knwo what az its in | |
| 14:22:31 | sean-k-mooney | because that shoudl still be changable vai the api | |
| 14:22:43 | dansmith | well, I don't see "desired_az" as breaking the "knowing what az I'm in" rule | |
| 14:22:49 | dansmith | yeah of course | |
| 14:23:05 | sean-k-mooney | desired_az to me woudl be like the initial cpu allocation ratio | |
| 14:23:10 | dansmith | just like our default allocation ratios - they're used if we have to set a default, otherwise it's what is in placement that matters | |
| 14:23:13 | dansmith | yep | |
| 14:23:14 | sean-k-mooney | rigth somethign that is used once and only once | |
| 14:23:17 | dansmith | exactly | |
| 14:23:24 | bauzas | fwiw, AZs are already a complicated matter to understand | |
| 14:23:36 | sahid | dansmith: yes our biggest concern was the az showed | |
| 14:23:38 | bauzas | do we really want to introduce another concept ? | |
| 14:23:51 | dansmith | bauzas: what other concept? | |
| 14:24:14 | bauzas | 'I, as compute, will be part of an AZ eventually' | |
| 14:24:33 | dansmith | well, that's true regardless of what we do here :) | |
| 14:24:33 | sean-k-mooney | that a concept for the installer/deploer | |
| 14:24:36 | bauzas | reconciliation, if you prefer | |
| 14:24:42 | sean-k-mooney | not really for the end users and only partly for the admin | |
| 14:25:09 | bauzas | dansmith: see, sahid's problem is the user-facing AZ list problem | |
| 14:25:30 | bauzas | not the day-2 operational problem of adding a host to an AZ | |
| 14:25:35 | dansmith | bauzas: I've said several times, I get that.. I think there are two benefits to a solution here, his being one | |
| 14:25:46 | sean-k-mooney | im still not sure if that can be adress by settign default_avaialblity_zone=internal | |
| 14:26:04 | bauzas | sean-k-mooney: internal has a very different meaning today | |
| 14:26:09 | sean-k-mooney | https://docs.openstack.org/nova/latest/configuration/config.html#DEFAULT.internal_service_availability_zone | |
| 14:26:19 | sean-k-mooney | bauzas: it does but its not show to non admins | |
| 14:26:27 | sean-k-mooney | and sahid does not want this to be show | |
| 14:26:32 | bauzas | internal skips all computes | |
| 14:26:45 | sean-k-mooney | i know | |
| 14:27:01 | bauzas | again, this show problem is resolvable by default_availability_zone | |
| 14:27:01 | sean-k-mooney | but do we have somethign to prevent you setting https://docs.openstack.org/nova/latest/configuration/config.html#DEFAULT.default_availability_zone to internal | |
| 14:27:11 | bauzas | ah, no | |
| 14:27:20 | bauzas | but I can quickly check | |
| 14:28:13 | sean-k-mooney | i would be ok with extending the concep of the intenal az to have compute in it if you set default_availability_zone=internal and documenting this means they cannot be used for workloads | |
| 14:28:40 | sean-k-mooney | i.e. make it a parkign ground for newly added compute serivce that the operator is yet to assign to a real az | |
| 14:28:53 | sean-k-mooney | i dont think that woudl break any exising usecases | |
| 14:29:21 | bauzas | this could break the filter | |
| 14:29:23 | bauzas | https://github.com/openstack/nova/blob/033af941792a9ae510c8d6b2cc318f062f0e1c66/nova/scheduler/filters/availability_zone_filter.py#L68 | |