Earlier  
Posted Nick Remark
#openstack-nova - 2023-05-03
14:17:26 bauzas dansmith: I'm confused, everything is API-driven
14:17:35 dansmith sahid: is the confusion of showing the wrong AZ the only thing you care about?
14:17:47 dansmith bauzas: adding a new compute is not api-driven
14:17:55 dansmith bauzas: the exception to that is needing to put it into the right az
14:18:43 sean-k-mooney bauzas: expanding on ^ you can delete compute services from the api, you can only create them via rpc via the conductor
14:19:48 sean-k-mooney bauzas: we do not have a create endpoint https://docs.openstack.org/api-ref/compute/#compute-services-os-services
14:20:08 bauzas OK I admit the feature gap
14:20:16 sean-k-mooney dansmith: addign a create endpoint woudl also work i guess. when the compute starts it would find the existing record
14:20:40 sean-k-mooney and ocne the record is created you could perhaps add the servicew to an aggreate
14:20:41 dansmith sean-k-mooney: to do that we'd have to expose the concept of a cell to the API which I'm -3 on :)
14:20:41 bauzas but I'm very reluctant to make add_host_to_agg() less strict
14:20:59 sean-k-mooney oh ya we would
14:21:00 dansmith making this more manual is also not something I'm in favor of
14:21:05 sean-k-mooney case this is in the cell db
14:21:12 sean-k-mooney so ya thats not an option relaly
14:21:21 dansmith and pre-creating hosts makes that more manual, and much more likely to be wrong, given what we know about hostnames :)
14:21:34 sean-k-mooney very true
14:21:35 dansmith I'd like to make this more automatic
14:21:49 dansmith which may or may not be what sahid really wants, but making it more automatic will *also* solve sahid's concern
14:21:50 bauzas without any upcalls for sure
14:21:55 dansmith yep
14:22:12 bauzas which is quite a chicken-and-egg problem
14:22:14 sean-k-mooney the thing is we would like this to be more atomatic while also not havign upcalls or having the comptue specificly knwo what az its in
14:22:31 sean-k-mooney because that shoudl still be changable vai the api
14:22:43 dansmith well, I don't see "desired_az" as breaking the "knowing what az I'm in" rule
14:22:49 dansmith yeah of course
14:23:05 sean-k-mooney desired_az to me woudl be like the initial cpu allocation ratio
14:23:10 dansmith just like our default allocation ratios - they're used if we have to set a default, otherwise it's what is in placement that matters
14:23:13 dansmith yep
14:23:14 sean-k-mooney rigth somethign that is used once and only once
14:23:17 dansmith exactly
14:23:24 bauzas fwiw, AZs are already a complicated matter to understand
14:23:36 sahid dansmith: yes our biggest concern was the az showed
14:23:38 bauzas do we really want to introduce another concept ?
14:23:51 dansmith bauzas: what other concept?
14:24:14 bauzas 'I, as compute, will be part of an AZ eventually'
14:24:33 dansmith well, that's true regardless of what we do here :)
14:24:33 sean-k-mooney that a concept for the installer/deploer
14:24:36 bauzas reconciliation, if you prefer
14:24:42 sean-k-mooney not really for the end users and only partly for the admin
14:25:09 bauzas dansmith: see, sahid's problem is the user-facing AZ list problem
14:25:30 bauzas not the day-2 operational problem of adding a host to an AZ
14:25:35 dansmith bauzas: I've said several times, I get that.. I think there are two benefits to a solution here, his being one
14:25:46 sean-k-mooney im still not sure if that can be adress by settign default_avaialblity_zone=internal
14:26:04 bauzas sean-k-mooney: internal has a very different meaning today
14:26:09 sean-k-mooney https://docs.openstack.org/nova/latest/configuration/config.html#DEFAULT.internal_service_availability_zone
14:26:19 sean-k-mooney bauzas: it does but its not show to non admins
14:26:27 sean-k-mooney and sahid does not want this to be show
14:26:32 bauzas internal skips all computes
14:26:45 sean-k-mooney i know
14:27:01 bauzas again, this show problem is resolvable by default_availability_zone
14:27:01 sean-k-mooney but do we have somethign to prevent you setting https://docs.openstack.org/nova/latest/configuration/config.html#DEFAULT.default_availability_zone to internal
14:27:11 bauzas ah, no
14:27:20 bauzas but I can quickly check
14:28:13 sean-k-mooney i would be ok with extending the concep of the intenal az to have compute in it if you set default_availability_zone=internal and documenting this means they cannot be used for workloads
14:28:40 sean-k-mooney i.e. make it a parkign ground for newly added compute serivce that the operator is yet to assign to a real az
14:28:53 sean-k-mooney i dont think that woudl break any exising usecases
14:29:21 bauzas this could break the filter
14:29:23 bauzas https://github.com/openstack/nova/blob/033af941792a9ae510c8d6b2cc318f062f0e1c66/nova/scheduler/filters/availability_zone_filter.py#L68
14:29:50 sahid bauzas: if user can see that, they can try to deploy instance with this az which is not what we want as-well
14:30:06 sean-k-mooney not if we reject all api request with az=Config.internal_service_availability_zone
14:30:31 sean-k-mooney sahid: i belive they cannot see it unless they are an admin
14:30:36 sahid (i mean in an ideal case)
14:31:10 bauzas sahid: again, speaking of examples
14:31:25 bauzas sahid: if you already have AZ1, AZ2 and AZ2 that are meaningful
14:31:37 bauzas all those three are shown by nova az-list
14:31:49 bauzas and now you do care of not showing 'nova' in that list
14:32:12 bauzas what I'm saying is that then change 'default_az' to any of the three, and job is done
14:35:50 sahid ok, let's use this way
15:06:11 opendevreview Sylvain Bauza proposed openstack/nova master: Fix get_segments_id with subnets without segment_id https://review.opendev.org/c/openstack/nova/+/882160
15:12:57 dansmith eharney: if you're good with this, it could use a +W as everything it depends on is in the gate now: https://review.opendev.org/c/openstack/cinder-tempest-plugin/+/881764
15:18:58 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
15:21:41 dansmith man, the nova jobs sure have gotten big again
15:21:50 dansmith we're running a lot of jobs
15:28:04 auniyal dansmith, yes its takes around 2 hour for a patch to get verified from tox
15:28:24 auniyal zuul
15:32:20 dansmith auniyal: that's generally a function of the slowest job, not the number of jobs
16:50:18 dansmith gouthamr: the tempest stuff is all landed.. if you can +W the cinder-tempest-change now, it can go in and then the devstack plugin change is all that's left
16:56:22 dansmith eharney: good catch I guess... I don't know if it's okay to just bump the requirement or not
16:56:25 dansmith gmann: https://review.opendev.org/c/openstack/cinder-tempest-plugin/+/881764?tab=comments
16:57:13 dansmith 27 was wallaby and 30 was yoga
16:57:36 dansmith so upstream I would think we're good bumping the version but I dunno if anything installs this from master in older versions...
17:25:15 gmann dansmith: give me 2 min, will check
17:25:36 dansmith gmann: ack thanks
17:34:49 gouthamr stable jobs install "master" version of the plugin (and tempest) unless explicitly pinned
17:35:57 dansmith gouthamr: but how far back?
17:36:19 dansmith we pin tempest at some point I think when the stable jobs get too old to support master tempest
17:37:08 dansmith all the release jobs back to xena are passing
17:37:39 gouthamr ah good; the first pin i see is in stable/wallaby: https://github.com/openstack/cinder/blob/stable/wallaby/.zuul.yaml#L291-L292
17:37:47 dansmith and I know they're using master tempest because they were honoring the depends-on when I was working on those patches (and breaking things)
17:38:07 dansmith gouthamr: ack, and since the xena job on c-t-p works, we should be good then right?
17:38:45 gouthamr i think so dansmith; tosky and gmann would be experts in this area
17:38:59 dansmith ack, well, we'll see what gmann has to say
17:39:43 gouthamr for correctness, we need the latest version of tempest - and afawct, that's working as expected..
17:40:06 dansmith yeah, so I'm not sure what the tempest requirement would be, if we're actually using master tempest
17:40:11 gouthamr we just hope no-one outside of our gates is pinning tempest for whatever reason but using a newer version of ctp.. (a weirdness i've seen albeit temporarily in some rdo jobs)
17:40:50 gmann dansmith: gouthamr tempest is pinned till wallaby as stable/xena still not in EM https://review.opendev.org/c/openstack/releases/+/881254
17:41:26 dansmith gmann: right, but xena is good according to this: https://review.opendev.org/c/openstack/cinder-tempest-plugin/+/881764?tab=change-view-tab-header-zuul-results-summary
17:42:16 gmann dansmith: ok, which one is failing
17:42:33 dansmith gmann: nothing is failing

Earlier   Later