Earlier  
Posted Nick Remark
#openstack-nova - 2020-03-25
17:06:53 melwitt kashyap: would you mind revisiting the aarch64 patch, it's been updated ^
17:09:32 openstackgerrit Lee Yarwood proposed openstack/nova master: WIP libvirt: Break up get_disk_mapping within blockinfo https://review.opendev.org/714962
17:12:05 sean-k-mooney alex_xu: dansmith gibi so just did a evacuate test with the cyborg fake driver. http://paste.openstack.org/show/791153/
17:12:06 openstackgerrit Stephen Finucane proposed openstack/nova master: libvirt: Add typing information https://review.opendev.org/714695
17:12:06 openstackgerrit Stephen Finucane proposed openstack/nova master: hardware: Update and correct typing information https://review.opendev.org/714694
17:12:07 openstackgerrit Stephen Finucane proposed openstack/nova master: objects: Replace 'cpu_pinning_requested' helper https://review.opendev.org/714697
17:12:07 openstackgerrit Stephen Finucane proposed openstack/nova master: tests: Split instance NUMA object tests https://review.opendev.org/714696
17:12:08 openstackgerrit Stephen Finucane proposed openstack/nova master: hardware: Remove handling of pre-Train compute nodes https://review.opendev.org/714699
17:12:08 openstackgerrit Stephen Finucane proposed openstack/nova master: hardware: Don't consider overhead CPUs for unpinned instances https://review.opendev.org/714698
17:12:09 openstackgerrit Stephen Finucane proposed openstack/nova master: hardware: Tweak the 'cpu_realtime_mask' handling slightly https://review.opendev.org/461456
17:12:09 openstackgerrit Stephen Finucane proposed openstack/nova master: hardware: Add validation for 'cpu_realtime_mask' https://review.opendev.org/468203
17:12:10 openstackgerrit Stephen Finucane proposed openstack/nova master: hardware: Invert order of NUMA topology generation https://review.opendev.org/714701
17:12:10 openstackgerrit Stephen Finucane proposed openstack/nova master: hardware: Rework 'get_realtime_constraint' https://review.opendev.org/714700
17:12:49 sean-k-mooney alex_xu: dansmith gibi we can evacuate but it does not create allocation for the fpga
17:15:03 sean-k-mooney the arqs are also not updated http://paste.openstack.org/show/791154/
17:15:34 sean-k-mooney ill update the block operation patch review with that info but currently we cannot evacuate properly.
17:16:08 lyarwood stephenfin: https://review.opendev.org/#/c/696834/ - not sure if you're still here but this should be ready now.
17:18:50 stephenfin lyarwood: done
17:19:30 lyarwood stephenfin: thanks
17:20:44 openstackgerrit John Garbutt proposed openstack/nova master: WIP: Enforce resource limits using oslo.limit https://review.opendev.org/615180
17:21:46 openstackgerrit John Garbutt proposed openstack/nova master: Prevent compute manager freeze when greenpool is full https://review.opendev.org/575034
17:27:24 openstackgerrit melanie witt proposed openstack/nova stable/train: Add config option for neutron client retries https://review.opendev.org/715010
17:30:15 openstackgerrit John Garbutt proposed openstack/nova master: Prevent compute manager freeze when greenpool is full https://review.opendev.org/575034
17:31:54 melwitt lyarwood: I dunno if you saw my comment on this one https://review.opendev.org/708030 IIUC this is an option you're thinking to keep indefinitely, if so, it shouldn't go under [workarounds] as they're things intended to be temporary and removed
17:44:03 kashyap melwitt: Hiya; will look at the AArch64 thing tom. in the AM. (Aside: just to keep you posted, I'm off from tomm. evening until 31)
17:44:15 kashyap (s/31/31st-Mar/)
17:45:57 kashyap Actually, looking now
17:46:24 melwitt cool thanks!
17:47:03 kashyap melwitt: Okay, they went with the upstream QEMU AArch64 recomm. of model 'max'. Cool
17:52:09 openstackgerrit John Garbutt proposed openstack/nova master: WIP: Enforce resource limits using oslo.limit https://review.opendev.org/615180
18:05:08 kashyap stephenfin: melwitt: The release note contains a lot of not useful info, which will only confuse: https://review.opendev.org/#/c/709494/20
18:05:22 kashyap stephenfin: melwitt: I suggested whittling it down to a couple of sentences. Hope that looks okay
18:06:07 stephenfin kashyap: yeah, I was iffy on that too but figured it was good enough. Now that there's two of us...
18:06:13 kashyap Maybe whoever is merging it can amend it? If it's not urgent, perhaps Kevin could respoin
18:06:51 kashyap stephenfin: Hehe, much of it is verbatim from a review comment I made; looks odd to have "stream of consciounsess" as a release note ;-)
18:07:55 kashyap Err, I myself made a grammar error; /me goes to fix
18:09:22 kashyap Alright; /me goes to make some dinner
18:57:49 openstackgerrit Ghanshyam Mann proposed openstack/nova master: Add test coverage of existing hypervisors policies https://review.opendev.org/715029
19:10:23 openstackgerrit Ghanshyam Mann proposed openstack/nova master: Introduce scope_types in os-hypervisors https://review.opendev.org/715036
19:14:31 lyarwood melwitt: yeah sorry was working my way down to these changes this week
19:14:36 melwitt dansmith: are you aware that in a vanilla devstack with one cell, we are getting [workarounds]disable_group_policy_check_upcall = True ? this is new to me
19:14:56 lyarwood melwitt: I'll respin and/or update in the morning.
19:14:59 mriedem it's intentional because of superconductor mode
19:15:02 mriedem melwitt: ^
19:15:13 dansmith yeah, what mriedem said
19:15:14 mriedem otherwise affinity tests will below up
19:15:28 dansmith so people either have to disable affinity or enable that workaround
19:15:29 mriedem there are a few things disabled by default like that in devstack
19:15:31 melwitt lyarwood: ok, np at all. just wanted to make sure in case you didn't see
19:15:33 dansmith because we still don't have affinity in placement
19:16:14 mriedem https://docs.openstack.org/nova/latest/user/cellsv2-layout.html#operations-requiring-upcalls
19:16:40 mriedem anything not marked complete in that list probably has some kind of flag to disable it in devstack
19:17:31 mriedem and we don't test cross_az_attach=false in the gate anywhere so...that just flies under the radar
19:17:35 melwitt mriedem, dansmith: thanks. yeah, I see that the logic is based on whether we have a superconductor going or not. just trying to work out whether we want or if there's way to set it False if we know everything's on the same MQ. context, we got a regression reported https://bugs.launchpad.net/nova/+bug/1863190 that's looking not like a regression since I saw [workarounds]disable_group_policy_check_upcall = True in the config
19:17:36 openstack Launchpad bug 1863190 in OpenStack Compute (nova) "Server group anti-affinity no longer works" [Undecided,New]
19:18:26 mriedem i think there is at least one known latent multi-cell bug with how (anti-)affinity works, but i'm fuzzy on the details
19:18:28 melwitt two parallel anti-affinity requests seeming to violate policy in the single MQ deployment
19:18:33 mriedem if you don't have multiple cells then i guess that doesn't apply
19:19:01 mriedem if you have single cell and are support anti-affinity then you need the late check enabled in the compute
19:19:02 melwitt yeah if you are multi MQ then you can't get anti-affinity if the requests hit at the same time
19:19:11 melwitt right
19:19:37 mriedem [workarounds]disable_group_policy_check_upcall = True means you're opting into the wildness
19:19:43 melwitt I was just thinking you'd think our default devstack with one MQ should set it False
19:19:50 mriedem it's false by default
19:20:12 melwitt it is, but something in the default devstack logic is setting it True
19:20:32 melwitt that is, I cloned the devstack repo and brought up a vanilla devstack and I'm getting it set to True
19:20:34 mriedem yeah, because superconductor is the default mode
19:20:51 mriedem superconductor is the ideal mode for deploying nova, so we test in the gate with that by default
19:21:11 melwitt yeah. and I'm thinking this sort of "bug" will keep being reported occasionally bc people not realizing what devstack is doing
19:21:32 mriedem it's a known limitation, link them to the docs
19:22:02 mriedem tempest has (anti)affinity tests as well but i don't think they make parallel requests
19:22:09 mriedem because of this
19:22:29 melwitt yeah. they have to be very parallel too. the first time I tried to repro I didn't get it but after a few tries I got it
19:24:41 mriedem right, i remember poking around in those tempest tests awhile back related to all of this
19:26:33 mriedem oh also i think tempest tests the affinity stuff with 2 servers in the same create request - same (anti)affinity group, and that works b/c the scheduler knows about the decisions within the same request
19:26:55 melwitt ah, yeah
19:27:06 mriedem https://github.com/openstack/tempest/blob/7a588ded216f74ddd0015c3065d4fae10de2161f/tempest/api/compute/admin/test_servers_on_multinodes.py
19:27:46 mriedem and https://github.com/openstack/tempest/blob/f419f4d36fd0f99a9c53fe3a984d172b02e828c5/tempest/api/compute/servers/test_server_group.py
19:28:24 mriedem but of course you get nfv mano systems in the wild that are robots just firing off rapid requests
19:28:50 mriedem but those are likely single cell and shouldn't be disabling that late affinity check upcall :)
19:29:12 mriedem starlingx had patches for a lot of this server group stuff
19:29:20 melwitt yeah, I see
19:29:57 mriedem to try and mitigate some of it, but it wasn't all perfect either, e.g. locks within conductor but that would only lock *that* conductor (or scheduler) worker, not across all - unless you use an external locking mechanism, like a db or etcd or something
19:30:22 melwitt right
19:30:22 melwitt yeah, I know of one where they run with a single scheduler and serialize affinity requests
19:30:26 mriedem trying to make scheduling requests serialized for group based scheduling
19:30:35 mriedem for starlingx with like 1 node and 1 worker then it's probably fine
19:32:08 melwitt ok. I'll finish repro'ing the situation with the workaround set to False to make sure the late affinity check triggers, and write up something for the bug. and will link the doc
19:44:45 mriedem melwitt: if you're so inclined and it's something that keeps coming up it might be worth writing something up in the troubleshooting docs, e.g. why are my servers that are in x policy landing on the same/different hosts when they shouldn't?
19:44:53 mriedem and explain the parallel issue
19:45:24 mriedem i found it nice to write something up once and then just point people to that
19:45:41 mriedem https://docs.openstack.org/nova/latest/admin/support-compute.html
19:51:01 melwitt mriedem: yup I think that's a good idea. I'll do that
19:51:09 melwitt thanks for suggesting
19:55:47 mriedem my first few weeks on the new job were me in slack being like "why x? why y? where is z documented?" and then taking the answers and trying to document them to feel like i was useful
19:58:22 melwitt that's a good investment. every time I don't do that, I regret it later
19:59:31 melwitt and I usually forget because I have the memory recall of a hamster
20:37:49 melwitt johnsom: hey, finally got a chance to dig into the bug report you opened awhile back about anti-affinity, tl;dr is I don't find that there's been a regression. pls see my latest comment explaining https://bugs.launchpad.net/nova/+bug/1863190
20:37:50 openstack Launchpad bug 1863190 in OpenStack Compute (nova) "Server group anti-affinity no longer works" [Undecided,New]
20:38:46 johnsom melwitt Ok, thank you for having a look. I got some feedback that it changed around queens, but I didn't go back and confirm either way.

Earlier   Later