| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-05-03 | |||
| 13:42:00 | sean-k-mooney | dansmith: we could also avoid the upcall by having the nova comptue just call the nova api driectly | |
| 13:42:07 | sahid | this should resolve our case for sure and is a bit like what sean-k-mooney mentioned | |
| 13:42:31 | dansmith | sean-k-mooney: it needs admin credentials though which would be unfortunate I think | |
| 13:42:46 | sean-k-mooney | well at least teh service role | |
| 13:42:52 | sean-k-mooney | but ya | |
| 13:43:00 | dansmith | yeah, not full admin hopefully, but still | |
| 13:43:11 | sean-k-mooney | where where you thinking of stashing the desired az | |
| 13:43:34 | sean-k-mooney | i dont think we have a db table we coudl use for that in the cell db currenlty | |
| 13:43:34 | dansmith | we'd need to add a generic metadata blob to service | |
| 13:43:40 | dansmith | nope | |
| 13:43:58 | sean-k-mooney | thats also not partically nice but ok | |
| 13:44:24 | dansmith | not nice because it requires db changes? or not nice for other reasons? | |
| 13:44:33 | sean-k-mooney | the db changes | |
| 13:44:38 | sean-k-mooney | if its just for this | |
| 13:44:48 | sean-k-mooney | if it was like instance_system_metadta | |
| 13:44:55 | sean-k-mooney | and we had other usecases i woudl care less | |
| 13:45:06 | dansmith | yeah, well, I mean.. something will require changes.. but I figure we might be able to use it for other things too | |
| 13:45:30 | sean-k-mooney | so you prefer doign ti via discover host to limit the creds | |
| 13:45:53 | dansmith | well, we already have a discovery process to do very similar things | |
| 13:46:09 | dansmith | lemme finish listening to this meeting for the next 15 minutes and then I can think a little more | |
| 13:51:01 | sahid | this should resolve our case for sure and is a bit like what sean-k-mooney mentioned | |
| 13:51:07 | sahid | oops sorry | |
| 13:56:09 | sahid | i've create bug report here, https://bugs.launchpad.net/nova/+bug/2018398 in case that we consider it's as valid and we want to do something to improve users experience :-) | |
| 13:57:02 | dansmith | sahid: s/users/operators/ so we're clear which "users" we're talking about | |
| 13:57:38 | dansmith | but yeah, I think this would be a nice behavior for operators, as long as we can do it within the other constraints and design points we have | |
| 13:57:52 | bauzas | sahid: so thanks for the bug report, now I better understand your concern | |
| 13:58:02 | bauzas | lemme check one thing tho | |
| 13:59:18 | bauzas | sahid: so, say your operator defined two AZs : AZ1 and AZ2 | |
| 13:59:32 | bauzas | hostA is in AZ1, hostB in AZ2 | |
| 13:59:34 | sahid | dansmith: wait, perhaps there is something that we are not doing correctly, but it's really our users which see those AZs in our interfaces | |
| 13:59:42 | bauzas | now, he gonna add hostC | |
| 14:00:19 | bauzas | once he registers hostC, then it appears a third AZ in the AZ list which is 'default_availability_zone' ie. 'nova' | |
| 14:00:31 | bauzas | I get the problem | |
| 14:00:32 | sahid | when they create an instance they can select the AZ | |
| 14:00:36 | bauzas | nown my question | |
| 14:00:51 | sahid | bauzas: yes rights | |
| 14:00:55 | bauzas | why isn't the default availabilty zone be either AZ1 or AZ2 ? | |
| 14:01:13 | bauzas | we just use that value to fake AZs | |
| 14:01:26 | dansmith | sahid: I know users see AZs, but users will not see this compute node feature | |
| 14:01:29 | bauzas | but this is not meaning that instances are automatically scheduled to that ZA | |
| 14:01:32 | bauzas | AZ* | |
| 14:01:36 | sahid | dansmith: ack | |
| 14:01:56 | bauzas | the default value for sending an instance to an AZ is default_schedule_zone | |
| 14:02:03 | bauzas | which is None by default | |
| 14:02:37 | dansmith | bauzas: oh right, I always forget about this difference | |
| 14:02:42 | dansmith | there are two conf knobs right? | |
| 14:03:52 | bauzas | yes | |
| 14:04:36 | bauzas | https://docs.openstack.org/nova/latest/configuration/config.html#DEFAULT.default_schedule_zone is for pushing new instances to a specific AZ automatically | |
| 14:05:07 | bauzas | https://docs.openstack.org/nova/latest/configuration/config.html#DEFAULT.default_availability_zone is just for showing a AZ if none exists | |
| 14:05:35 | bauzas | in a context where some AZs already exist, the solution is just to set the latter option to an existing AZ | |
| 14:05:38 | dansmith | bauzas: but it doesn't actually put the compute there right? | |
| 14:05:45 | bauzas | no | |
| 14:05:49 | dansmith | it's like faked in the api? | |
| 14:05:53 | sean-k-mooney | yep | |
| 14:05:54 | bauzas | " This option determines the default availability zone for ‘nova-compute’ services, which will be used if the service(s) do not belong to aggregates with availability zone metadata." | |
| 14:06:02 | sean-k-mooney | its not used by the compute service at all | |
| 14:06:40 | sean-k-mooney | you could have discover host use it if its set to a real az | |
| 14:06:53 | dansmith | yeah, so now does that help? I mean, it might help the "don't expose some other AZ for a window after adding the compute" but it doesn't help the mechanical operations of just dealing with the compute itself while adding | |
| 14:07:11 | bauzas | dansmith: there are two different concepts that are hard to understand | |
| 14:07:12 | dansmith | sean-k-mooney: but then discover would add all computes to one az, not the one the compute should be in | |
| 14:07:17 | bauzas | there is a the AZ API | |
| 14:07:28 | sean-k-mooney | dansmith: ya its not greate to use that directly | |
| 14:07:38 | bauzas | which is just showing the list of aggregates with a AZ metadata + the default az value | |
| 14:07:50 | bauzas | and then, the second concept is the scheduling thing | |
| 14:08:05 | dansmith | yeah, I understand the default_schedule_zone part | |
| 14:08:30 | bauzas | dansmith: if you read https://bugs.launchpad.net/nova/+bug/2018398 sahid is concerned by the fact 'nova' is presented to endusers when they do nova availability-zone-list | |
| 14:08:33 | sean-k-mooney | what happens today if you set default_avaiablity_zone=internal | |
| 14:08:39 | dansmith | bauzas: right | |
| 14:08:47 | sean-k-mooney | internal is not show to endusers unless they are admin | |
| 14:09:03 | bauzas | dansmith: I'm just saying to modify this value to an existing AZ value, eg. AZ1 | |
| 14:09:05 | dansmith | bauzas: although I'm not sure how that happens even, because the compute isn't actually in that az | |
| 14:09:18 | bauzas | dansmith: that option is purely API-based | |
| 14:09:20 | sean-k-mooney | if default_avaiablity_zone=internal is accpeted today | |
| 14:09:26 | dansmith | bauzas: right, that helps not expose 'nova' for a minute while you fix it | |
| 14:09:29 | sean-k-mooney | that woudl hide the compute until the admin added them to a real az | |
| 14:09:34 | bauzas | correct ^ | |
| 14:09:37 | dansmith | bauzas: but that's only one aspect of the problem (IMHO) | |
| 14:09:49 | bauzas | dansmith: well, lemme clarify then | |
| 14:09:55 | dansmith | bauzas: the second aspect is that you add a compute, you wait for it to register, you wait for it to be discovered, then you go and add it to the right az | |
| 14:10:16 | dansmith | that's stuff that a deployment tool has to do a bunch of polling and looping over, but it really shouldn't | |
| 14:10:40 | sean-k-mooney | dansmith: if we remvoed the exists check in the host aggrate add | |
| 14:10:45 | dansmith | if it could say "this should be in AZ foo" then all that stuff can happen automatically and without a bunch of slow steps in the process | |
| 14:10:47 | bauzas | dansmith: technically you don't add a host to an az, but to an aggregate | |
| 14:10:55 | sean-k-mooney | and we leveraged the fact we can right a uuid for the comptue to use | |
| 14:11:02 | dansmith | sean-k-mooney: yeah but then people can typo things when they add them, so I don't love that | |
| 14:11:16 | dansmith | bauzas: I know :) | |
| 14:11:20 | bauzas | dansmith: and an operator can add a host to an aggregate as soon as it's registered in the services table | |
| 14:11:20 | sean-k-mooney | i could see the workflow being the installer generate a uuid for the host, then adds it to the righ host aggrate and then deploys the host | |
| 14:11:41 | sean-k-mooney | dansmith: ya that one of the reasons we dont allow it today | |
| 14:11:44 | dansmith | sean-k-mooney: but you don't add compute nodes to aggregates, you add services right? | |
| 14:11:56 | sean-k-mooney | ah you are right its the service | |
| 14:11:56 | bauzas | dansmith: typos can't append, we query the record based on the service name | |
| 14:12:04 | sean-k-mooney | which is why this does nto really work with ironic today | |
| 14:12:05 | bauzas | happen* | |
| 14:12:21 | dansmith | bauzas: sean-k-mooney is suggesting we drop that requirement | |
| 14:12:33 | sean-k-mooney | dansmith: well sahid was previously | |
| 14:12:35 | bauzas | again, why ? | |
| 14:12:53 | sean-k-mooney | bauzas: to allwo you to pregreister the service to host aggrate | |
| 14:12:56 | sean-k-mooney | before its created | |
| 14:13:00 | bauzas | can't a script loop over a service list and do the aggregate add after ? | |