| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-02-08 | |||
| 20:01:35 | artom | I know it's an RPC cast and everything, so compute needs to be running to receive it | |
| 20:02:12 | mriedem | yes | |
| 20:02:15 | mriedem | correct i mean | |
| 20:02:24 | artom | But just sanity-checking, you can't mark an interface as detached and then later unplug the vif when compute comes back, right? | |
| 20:02:24 | melwitt | yeah. you could at best tell neutron to do things with the port | |
| 20:02:40 | dansmith | artom: not currently, but also, | |
| 20:02:46 | dansmith | consider if compute manager is all that is down, | |
| 20:02:54 | dansmith | but the interface and address are still being used on the data plane | |
| 20:02:59 | dansmith | that would be like bad and stuff to re-assign it | |
| 20:03:17 | artom | dansmith, ah, indeed. | |
| 20:03:29 | artom | So not only we don't do it, but we don't even want to do it | |
| 20:03:54 | artom | Thanks dudes :) | |
| 20:03:58 | artom | (And Mel) | |
| 20:03:59 | artom | ;) | |
| 20:04:22 | melwitt | :) | |
| 20:07:12 | openstackgerrit | Ameed Ashour proposed openstack/osc-placement master: Change documentation theme https://review.openstack.org/542378 | |
| 20:10:27 | openstackgerrit | Dan Smith proposed openstack/nova master: Compute RPC client bump to 5.0 https://review.openstack.org/541035 | |
| 20:10:28 | openstackgerrit | Dan Smith proposed openstack/nova master: Clean up reservations in migrate_task call path https://review.openstack.org/542409 | |
| 20:10:49 | dansmith | artom: not without a lot more infrastructure around it I would say | |
| 20:10:50 | dansmith | mriedem: ^ as promised with cleanup | |
| 20:16:14 | mriedem | dansmith: https://review.openstack.org/#/c/541035/5/nova/compute/rpcapi.py@752 | |
| 20:16:21 | mriedem | my point there was the same as the reservations things you removed | |
| 20:17:22 | dansmith | mriedem: replied | |
| 20:17:47 | mriedem | sure, my point is, pass migration_id=None | |
| 20:17:53 | dansmith | but that wasn't valid | |
| 20:17:58 | dansmith | reservations=None was | |
| 20:18:08 | dansmith | something else will try to use migration_id as an integer and fail | |
| 20:18:47 | dansmith | oh, it was defaulted to none | |
| 20:19:21 | dansmith | FFS, fine.. we have migration still passed here and don't for reservations, so it seems like the most compatible to just do the right thing but whatever | |
| 20:19:38 | mriedem | long ago it didn't have a default of None https://review.openstack.org/#/c/287997/19/nova/compute/manager.py@5167 | |
| 20:20:10 | dansmith | ah, that was tdurakov's mistake yeah, | |
| 20:20:13 | mriedem | we could also just remove that later in cleanup | |
| 20:20:20 | dansmith | it's not compatible that way | |
| 20:20:46 | dansmith | so passing None there would break that older code if we were to interact with ti | |
| 20:20:53 | dansmith | mriedem: all the 4.x stuff goes away in rocky anyway | |
| 20:21:04 | dansmith | so yes, the same cleanup that removes all this would remove that anyway | |
| 20:21:13 | openstackgerrit | melanie witt proposed openstack/nova master: Make scheduler.utils.setup_instance_group query all cells https://review.openstack.org/540258 | |
| 20:21:19 | mriedem | +2 | |
| 20:21:22 | mriedem | onto the cleanup patch | |
| 20:22:16 | dansmith | I havent' written the "drop 4.x" patch yet, but can get on that once we start to slow down | |
| 20:22:52 | mriedem | +2s all the way | |
| 20:22:56 | mriedem | time to find another core | |
| 20:23:00 | mriedem | they tend to hide | |
| 20:26:08 | openstackgerrit | Merged openstack/nova master: Workaround glanceclient bug when CONF.glance.api_servers not set https://review.openstack.org/541008 | |
| 20:26:19 | openstackgerrit | Merged openstack/nova master: Add a prelude release note for the 17.0.0 Queens GA https://review.openstack.org/541495 | |
| 20:38:57 | dansmith | gah pep8 | |
| 20:40:53 | openstackgerrit | Dan Smith proposed openstack/nova master: Clean up reservations in migrate_task call path https://review.openstack.org/542409 | |
| 20:42:46 | cdent | mriedem, dansmith: wrote up some of the scale stuff I was playing with last night (and today too): https://anticdent.org/placement-scale-fun.html (efried, jaypipes might be of interest to you too). Some conclusions/observations towards the end. | |
| 20:45:14 | mriedem | ok | |
| 20:45:16 | mriedem | thanks | |
| 20:46:33 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Encode libvirt domain XML in UTF-8 https://review.openstack.org/522161 | |
| 20:47:12 | openstackgerrit | melanie witt proposed openstack/nova master: Refine waiting for vif plug events during _hard_reboot https://review.openstack.org/541442 | |
| 20:47:15 | dansmith | cdent: conductors do more than boot things, so conductors=schedulers isn't right | |
| 20:47:25 | dansmith | you'll always need more conductors than schedulers | |
| 20:47:53 | melwitt | gdi I forgot another thing I wanted to update | |
| 20:47:54 | cdent | dansmith: the point I was trying to make there was that _if_ you have eight conductors and you slam the api, then the conductor is going to slam the scheduler | |
| 20:47:59 | dansmith | you're measuring mostly boot load there which is why it seems like they're the same | |
| 20:48:05 | mriedem | dansmith: melwitt: need another +2 on https://review.openstack.org/#/c/522161/ - tempest regression test is passing with that, and it's clearly a regression in queens because the tempest test before i added the depends-on failed in master and passed on all the stable branches | |
| 20:48:30 | cdent | so at the starting point you're always going to overwhelm the scheduler processes | |
| 20:48:44 | cdent | with an otherwise initially idle conductor | |
| 20:48:50 | dansmith | cdent: also s/rcp timeout/rpc timeout/ | |
| 20:48:54 | cdent | thanks | |
| 20:49:26 | dansmith | cdent: also, the reasoning for the lack of workers= on the scheduler is that it used to be that >1 would just kill you because they'd stomp all over each other | |
| 20:49:34 | dansmith | with claims in the scheduling process, that stops being a problem | |
| 20:50:11 | cdent | yeah, bauzas and I were talking about that earlier today | |
| 20:50:26 | dansmith | ack | |
| 20:50:42 | dansmith | obviously the next thing would be to figure out what scheduler is doing the most of in there | |
| 20:51:13 | cdent | when I had 8 workers, the scheduler seemed to work just fine, it didn't show unreasonable suffering | |
| 20:52:28 | melwitt | mriedem: that's interesting, I had thought in the past we have encoded only the instance name in the XML to let non-ascii names work, not the entire XML. but I never knew if there's any harm in encoding the whole thing. must not be if that's what py3 does by default, right? | |
| 20:52:28 | dansmith | with one worker, all the load for scheduler work must be serialized, on a single core and in a single process | |
| 20:52:49 | dansmith | it will be penalized for that against anything else wanting to run | |
| 20:52:50 | dansmith | and on a system with everything running on it, it will get a smallish slice of the system, being CPU-bound | |
| 20:53:13 | openstackgerrit | melanie witt proposed openstack/nova master: Refine waiting for vif plug events during _hard_reboot https://review.openstack.org/541442 | |
| 20:53:14 | dansmith | everything else having multiple workers gives those things way more bandwidth and time slice, relative to scheduler, yet everything funnels through scheduler | |
| 20:53:30 | dansmith | which isn't necessarily how it would be on a real system with things separated, depending on the layout | |
| 20:54:01 | dansmith | but | |
| 20:54:16 | dansmith | on a multi-core system(i.e. anything from this decade) a single scheduler worker on the box makes no sense either | |
| 20:54:47 | dansmith | allowing workers= to be configured on the scheduler process is probably something we can/should do now | |
| 20:55:00 | dansmith | maybe default=1 for the moment, but get some feedback on that | |
| 20:55:03 | openstackgerrit | melanie witt proposed openstack/nova master: Make scheduler.utils.setup_instance_group query all cells https://review.openstack.org/540258 | |
| 20:55:32 | cdent | yeah, I tried to make it clear that I wasn't modeling a real world with what I was doing | |
| 20:55:52 | dansmith | no, I know, I'm just saying words | |
| 20:56:12 | cdent | it was pretty cool the way stuff mostly just worked | |
| 20:57:47 | mriedem | cdent: you could brush up https://review.openstack.org/#/c/159382/ and throw it into the env and see what difference it makes | |
| 20:58:34 | mriedem | heh i suppose i'd have to get jogo's -2 removed | |
| 20:58:35 | cdent | mriedem: I already did (my own version of) that. I had to go to eight workers to get things to not timeout | |
| 21:03:25 | cdent | mriedem: what I'm not clear on (yet) is what needs to be done to exercise any of the hoststate management concerns | |
| 21:03:46 | mriedem | in the scheduler? | |
| 21:03:54 | mriedem | using the default enabled filters is probably a good start | |
| 21:04:10 | dansmith | what does that mean? | |
| 21:04:15 | dansmith | hoststate management concerns | |
| 21:04:17 | mriedem | he disabled all of the filters | |
| 21:04:20 | mriedem | oh | |
| 21:04:31 | dansmith | oh, I missed that.. what is the scheduler doing then? | |
| 21:04:39 | mriedem | querying placement | |
| 21:04:43 | mriedem | fin | |
| 21:04:55 | dansmith | oh, yeah, I totally missed that part | |
| 21:05:16 | mriedem | cdent: another thing to consider is that all 75 fake computes are going to be rpc broadcasting the shit out of the scheduler(s) for all instance resource tracker type actions, and every update_available_resource periodic | |
| 21:05:16 | dansmith | that's not awesome, but maybe it's just the "8 people asking 1 person to do all the work" then | |
| 21:05:32 | dansmith | ...with no IO delays between them yeah | |