| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-05-03 | |||
| 14:01:29 | bauzas | but this is not meaning that instances are automatically scheduled to that ZA | |
| 14:01:32 | bauzas | AZ* | |
| 14:01:36 | sahid | dansmith: ack | |
| 14:01:56 | bauzas | the default value for sending an instance to an AZ is default_schedule_zone | |
| 14:02:03 | bauzas | which is None by default | |
| 14:02:37 | dansmith | bauzas: oh right, I always forget about this difference | |
| 14:02:42 | dansmith | there are two conf knobs right? | |
| 14:03:52 | bauzas | yes | |
| 14:04:36 | bauzas | https://docs.openstack.org/nova/latest/configuration/config.html#DEFAULT.default_schedule_zone is for pushing new instances to a specific AZ automatically | |
| 14:05:07 | bauzas | https://docs.openstack.org/nova/latest/configuration/config.html#DEFAULT.default_availability_zone is just for showing a AZ if none exists | |
| 14:05:35 | bauzas | in a context where some AZs already exist, the solution is just to set the latter option to an existing AZ | |
| 14:05:38 | dansmith | bauzas: but it doesn't actually put the compute there right? | |
| 14:05:45 | bauzas | no | |
| 14:05:49 | dansmith | it's like faked in the api? | |
| 14:05:53 | sean-k-mooney | yep | |
| 14:05:54 | bauzas | " This option determines the default availability zone for ‘nova-compute’ services, which will be used if the service(s) do not belong to aggregates with availability zone metadata." | |
| 14:06:02 | sean-k-mooney | its not used by the compute service at all | |
| 14:06:40 | sean-k-mooney | you could have discover host use it if its set to a real az | |
| 14:06:53 | dansmith | yeah, so now does that help? I mean, it might help the "don't expose some other AZ for a window after adding the compute" but it doesn't help the mechanical operations of just dealing with the compute itself while adding | |
| 14:07:11 | bauzas | dansmith: there are two different concepts that are hard to understand | |
| 14:07:12 | dansmith | sean-k-mooney: but then discover would add all computes to one az, not the one the compute should be in | |
| 14:07:17 | bauzas | there is a the AZ API | |
| 14:07:28 | sean-k-mooney | dansmith: ya its not greate to use that directly | |
| 14:07:38 | bauzas | which is just showing the list of aggregates with a AZ metadata + the default az value | |
| 14:07:50 | bauzas | and then, the second concept is the scheduling thing | |
| 14:08:05 | dansmith | yeah, I understand the default_schedule_zone part | |
| 14:08:30 | bauzas | dansmith: if you read https://bugs.launchpad.net/nova/+bug/2018398 sahid is concerned by the fact 'nova' is presented to endusers when they do nova availability-zone-list | |
| 14:08:33 | sean-k-mooney | what happens today if you set default_avaiablity_zone=internal | |
| 14:08:39 | dansmith | bauzas: right | |
| 14:08:47 | sean-k-mooney | internal is not show to endusers unless they are admin | |
| 14:09:03 | bauzas | dansmith: I'm just saying to modify this value to an existing AZ value, eg. AZ1 | |
| 14:09:05 | dansmith | bauzas: although I'm not sure how that happens even, because the compute isn't actually in that az | |
| 14:09:18 | bauzas | dansmith: that option is purely API-based | |
| 14:09:20 | sean-k-mooney | if default_avaiablity_zone=internal is accpeted today | |
| 14:09:26 | dansmith | bauzas: right, that helps not expose 'nova' for a minute while you fix it | |
| 14:09:29 | sean-k-mooney | that woudl hide the compute until the admin added them to a real az | |
| 14:09:34 | bauzas | correct ^ | |
| 14:09:37 | dansmith | bauzas: but that's only one aspect of the problem (IMHO) | |
| 14:09:49 | bauzas | dansmith: well, lemme clarify then | |
| 14:09:55 | dansmith | bauzas: the second aspect is that you add a compute, you wait for it to register, you wait for it to be discovered, then you go and add it to the right az | |
| 14:10:16 | dansmith | that's stuff that a deployment tool has to do a bunch of polling and looping over, but it really shouldn't | |
| 14:10:40 | sean-k-mooney | dansmith: if we remvoed the exists check in the host aggrate add | |
| 14:10:45 | dansmith | if it could say "this should be in AZ foo" then all that stuff can happen automatically and without a bunch of slow steps in the process | |
| 14:10:47 | bauzas | dansmith: technically you don't add a host to an az, but to an aggregate | |
| 14:10:55 | sean-k-mooney | and we leveraged the fact we can right a uuid for the comptue to use | |
| 14:11:02 | dansmith | sean-k-mooney: yeah but then people can typo things when they add them, so I don't love that | |
| 14:11:16 | dansmith | bauzas: I know :) | |
| 14:11:20 | sean-k-mooney | i could see the workflow being the installer generate a uuid for the host, then adds it to the righ host aggrate and then deploys the host | |
| 14:11:20 | bauzas | dansmith: and an operator can add a host to an aggregate as soon as it's registered in the services table | |
| 14:11:41 | sean-k-mooney | dansmith: ya that one of the reasons we dont allow it today | |
| 14:11:44 | dansmith | sean-k-mooney: but you don't add compute nodes to aggregates, you add services right? | |
| 14:11:56 | bauzas | dansmith: typos can't append, we query the record based on the service name | |
| 14:11:56 | sean-k-mooney | ah you are right its the service | |
| 14:12:04 | sean-k-mooney | which is why this does nto really work with ironic today | |
| 14:12:05 | bauzas | happen* | |
| 14:12:21 | dansmith | bauzas: sean-k-mooney is suggesting we drop that requirement | |
| 14:12:33 | sean-k-mooney | dansmith: well sahid was previously | |
| 14:12:35 | bauzas | again, why ? | |
| 14:12:53 | sean-k-mooney | bauzas: to allwo you to pregreister the service to host aggrate | |
| 14:12:56 | sean-k-mooney | before its created | |
| 14:13:00 | bauzas | can't a script loop over a service list and do the aggregate add after ? | |
| 14:13:16 | dansmith | I'm just thinking of how customers are sometimes afraid to run stack update on a running cloud, just like I'm afraid to re-run my own shoddy ansible sometimes :) | |
| 14:13:37 | dansmith | and how making the act of adding a new compute more of a compute-focused thing than a deployment-focused thing would likely be a welcome pattern for people | |
| 14:13:50 | bauzas | sean-k-mooney: we would soften the requirements for the very little benefit of pre-adding a compute to an aggregate | |
| 14:14:02 | sean-k-mooney | bauzas: you can but if you do that in cron and a vm is booted that request nova in that interval then it will be broken | |
| 14:14:06 | dansmith | bauzas: I agree, I don't like that idea | |
| 14:14:18 | sean-k-mooney | once an instance gets schdeuled to the host it cant move to a real az | |
| 14:14:46 | bauzas | I like the idea to only add something schedulable if that thing can actually really schedule my stuff | |
| 14:15:01 | sean-k-mooney | could we hide the nova az if all compute service in it were disabled | |
| 14:15:05 | bauzas | the second it got added to an aggregate, it can be scheduled | |
| 14:15:13 | bauzas | again, why ? | |
| 14:15:17 | sean-k-mooney | then you can set the config option to only register them as disabled | |
| 14:15:31 | bauzas | we would put a lot of preconditions and added complexity for a very little gain | |
| 14:15:48 | bauzas | and adding a compute isn't really a daily operation | |
| 14:15:57 | sean-k-mooney | bauzas: sure its little gain but its a ligitmate pain point for there end customer apperenlty | |
| 14:16:07 | sean-k-mooney | well it depend on your scale | |
| 14:16:34 | bauzas | sean-k-mooney: the bug report wasn't complaining about it | |
| 14:16:44 | dansmith | we already have a task that has to run for a new compute (discover) which can either be automated in scheduler, called via cron, or manually during a deployment.. its purpose is to do things in the api database, so it feels like a very natural thing to have that step also do this, which is placing the compute in the right grouping | |
| 14:16:51 | bauzas | sean-k-mooney: the bug report was about to prevent 'nova' to show off | |
| 14:16:58 | sean-k-mooney | bauzas: sahid said it caused confustion for ther custoemr. | |
| 14:17:26 | bauzas | dansmith: I'm confused, everything is API-driven | |
| 14:17:35 | dansmith | sahid: is the confusion of showing the wrong AZ the only thing you care about? | |
| 14:17:47 | dansmith | bauzas: adding a new compute is not api-driven | |
| 14:17:55 | dansmith | bauzas: the exception to that is needing to put it into the right az | |
| 14:18:43 | sean-k-mooney | bauzas: expanding on ^ you can delete compute services from the api, you can only create them via rpc via the conductor | |
| 14:19:48 | sean-k-mooney | bauzas: we do not have a create endpoint https://docs.openstack.org/api-ref/compute/#compute-services-os-services | |
| 14:20:08 | bauzas | OK I admit the feature gap | |
| 14:20:16 | sean-k-mooney | dansmith: addign a create endpoint woudl also work i guess. when the compute starts it would find the existing record | |
| 14:20:40 | sean-k-mooney | and ocne the record is created you could perhaps add the servicew to an aggreate | |
| 14:20:41 | bauzas | but I'm very reluctant to make add_host_to_agg() less strict | |
| 14:20:41 | dansmith | sean-k-mooney: to do that we'd have to expose the concept of a cell to the API which I'm -3 on :) | |
| 14:20:59 | sean-k-mooney | oh ya we would | |
| 14:21:00 | dansmith | making this more manual is also not something I'm in favor of | |
| 14:21:05 | sean-k-mooney | case this is in the cell db | |
| 14:21:12 | sean-k-mooney | so ya thats not an option relaly | |
| 14:21:21 | dansmith | and pre-creating hosts makes that more manual, and much more likely to be wrong, given what we know about hostnames :) | |
| 14:21:34 | sean-k-mooney | very true | |
| 14:21:35 | dansmith | I'd like to make this more automatic | |
| 14:21:49 | dansmith | which may or may not be what sahid really wants, but making it more automatic will *also* solve sahid's concern | |
| 14:21:50 | bauzas | without any upcalls for sure | |