Earlier  
Posted Nick Remark
#openstack-nova - 2018-09-27
14:47:53 bauzas mriedem: zactly
14:47:59 bauzas (16:46:11) bauzas: but I'm pretty sure we say it's INFO (and no ERROR or warning, because a capacity problem isn't a scheduling problem)
14:48:28 bauzas gosh, already 4:46pm here :(
14:48:48 mriedem so if we do persist the request spec requested_destination, i'm just not sure how it doesn't cause problems
14:49:52 mriedem maybe we just get lucky and don't call request_spec.save() on the dirty request spec?
14:50:16 mnaser would we be able to backport that log level change? it's kinda useful. if we can i guess i'll file a bug?
14:50:30 mriedem mnaser: sure
14:51:01 bauzas mriedem: https://github.com/openstack/nova/blob/master/nova/objects/request_spec.py#L29
14:51:37 mriedem bauzas: that doesn't really tell me anything
14:51:38 bauzas and shit, I threw my day on some internal bug and now I'm done, I have to go into a meeting
14:52:32 mriedem mnaser: this is your justification https://github.com/openstack/nova/blob/c6218428e9b29a2c52808ec7d27b4b21aadc0299/nova/filters.py#L130
14:52:46 mriedem b/c if we got allocation candidates, but the filters rejected all of them, we log something at INFO
14:53:25 mriedem http://logstash.openstack.org/#dashboard/file/logstash.json?query=message%3A%5C%22Filtering%20removed%20all%20hosts%20for%20the%20request%20with%5C%22%20AND%20tags%3A%5C%22screen-n-sch.txt%5C%22&from=7d
14:54:28 bauzas mriedem: I guess I have to doublecheck this spaghetti code
14:55:06 bauzas mriedem: but since requested_destination is only set on a live migration or an evacuation, I just wonder whether we .save() this
14:55:45 mriedem it's also set on resize
14:55:52 mriedem b/c you can pass a host on resize since queens
14:56:04 mriedem and on resize we persist the request spec with the new flavor before casting to compute
14:56:14 mriedem i'm pretty sure i raised this with takashi when he was writing that
14:56:54 bauzas ah
14:56:55 mriedem ah this is how he dealt with that https://github.com/openstack/nova/blob/master/nova/compute/api.py#L3505
14:56:56 bauzas ghood point
14:57:15 bauzas anyway, I need to jump on a call
14:57:17 mriedem but....that's likely not good enough if you resize to a specific host, and then live migrate without specifying a host...
14:57:51 bauzas mriedem: live migrate has the same logic IIRC
14:58:00 bauzas we null out the field
14:58:26 mriedem i don't see that happening
14:58:28 mriedem for live migrate
14:58:36 bauzas oh shit no you're right
14:58:43 bauzas bug bug bug
14:58:48 openstackgerrit Mohammed Naser proposed openstack/nova master: Use INFO for logging no allocation candidates https://review.openstack.org/605765
14:58:58 mnaser mriedem: bauzas ^
14:59:07 mnaser took me longer to come up with a decent commit message jeez
14:59:39 melwitt .
15:00:10 mriedem bauzas: i'll give myself a todo to write a regression test for this
15:00:23 bauzas mriedem: ack
15:01:06 mriedem mnaser: +2
15:01:54 openstackgerrit Christoph Manns proposed openstack/nova master: Fix stacktraces with redis caching backend https://review.openstack.org/605748
15:01:59 bauzas mnaser: +Wipped
15:02:15 bauzas mnaser: please make a cherry-pick for rocky
15:06:58 openstackgerrit Mohammed Naser proposed openstack/nova stable/rocky: Use INFO for logging no allocation candidates https://review.openstack.org/605771
15:07:09 mnaser bauzas: done
15:26:53 mnaser AggregateRamFilter is still relevant and working? i remember there was an aggregate filter that had a long ml discussion about how it wasnt really working?
15:28:14 mriedem mnaser: ask jaypipes re https://review.openstack.org/#/c/544683/ and https://review.openstack.org/#/c/552105/
15:38:06 melwitt mnaser: this is the situation http://lists.openstack.org/pipermail/openstack-dev/2018-January/126283.html and it's still the case now. those two specs ^ are what's needed to restore the ability to set allocation ratios per aggregate
15:41:14 openstackgerrit Balazs Gibizer proposed openstack/nova master: Ignore forcing of live migration for nested instance https://review.openstack.org/605785
15:41:33 gibi mriedem: my first stab for ignoring the force flag ^^
15:42:24 mgagne @mriedem: what's up with caching scheduler?
15:42:31 mriedem mgagne: are we ok to remove it now?
15:42:32 mriedem in stein
15:43:01 mnaser melwitt: i see, i think i might be looking at the wrong filter then
15:43:03 mriedem i.e. is heal_allocation sufficient for you right now to get upgraded to a FilterScheduler world
15:43:03 mgagne mriedem: I think we figured out it was ok after you wrote the allocation healing tool.
15:43:16 mgagne ++
15:43:17 mriedem mgagne: ok
15:43:18 mriedem thanks
15:43:24 mnaser maybe it was a weigher that had the flip of a switch
15:43:29 mnaser -1 or 1 to pack vs distribute
15:43:36 mriedem host_subset_size?
15:43:42 mnaser yeah it is gr
15:44:03 mnaser https://github.com/openstack/nova/blob/master/nova/scheduler/weights/ram.py i guess no way of having that per host or per aggregate or anything
15:44:30 mgagne mnaser: we have a private implementation of the RAMWeigher per aggregate
15:44:39 melwitt mgagne: while you're here, don't forget to re-propose https://review.openstack.org/312626 for stein. I was +2 on the implementation but we were at feature freeze at the time
15:45:06 mgagne melwitt: thanks for the follow up
15:45:39 mnaser mgagne: that'd be nice to have upstream i guess, assuming that's a possiblity
15:46:00 mgagne mnaser: sure, unfortunately, I'm not sure it's gonna play well with placement API :D
15:46:09 mgagne or maybe it's not related?
15:46:12 mnaser weighers are after placement
15:46:12 openstackgerrit Balazs Gibizer proposed openstack/nova master: Ignore forcing of live migration for nested instance https://review.openstack.org/605785
15:46:13 mnaser so it doesnt matter
15:46:18 mgagne mnaser: awesome
15:46:19 mnaser its just "which machine do i prefer"
15:46:39 mnaser but im not a nova dev but thats as far as i understand it
15:46:46 mnaser weighers run after placement allocations AND filters have ran
15:46:50 mgagne so you just need a high host_subset_size for weigher to get some hosts to choose from
15:47:37 mnaser yeah we have that bumped up. the silly idea is right now by default nova spreads vms which is ok but with large vms it becomes problematic
15:47:47 mnaser because we have enough capacity for them but just not in aggregate if that makes sense
15:47:54 mgagne hehe, I know =)
15:48:12 bauzas mriedem: you're my rebuild specialist, so lemme bug you about some silly question
15:48:19 mnaser but we had some really bad decision making with our first flavors so running it with -1 ram multiplier for * .. bad things happen
15:48:51 bauzas mriedem: as of today, do we rebuild by calling the scheduler or have we stopped this ? /me is confused by the number of bugs we had about
15:49:04 bauzas my brain sucks
15:50:50 openstackgerrit Elod Illes proposed openstack/nova master: Reject networks with QoS policy https://review.openstack.org/570079
15:51:40 mgagne mnaser: based on Newton (sorry): https://gist.github.com/mgagne/142e20e32049abd0cdf5d2da7e048608
15:51:42 melwitt mnaser: hm, supposed to pack by default, I thought
15:52:07 mnaser mnaser: by default the weighers are set to positive values so distribute
15:52:59 melwitt ok. in the past the scheduler used to pack by default, so I'm not sure when/how that changed
15:53:07 mgagne doesn't make much sense to pack since if you have lets say the openstack infra team spawning 120 VMs at the same time, you will overload the same hosts with requests.
15:53:35 melwitt not since claims in the scheduler, but in the past yeah
15:53:46 mgagne true
15:54:40 mgagne but still, maybe you don't want 6 images being downloaded for the first time on the same host at the same time. or other similar expansive operations.
15:54:48 mnaser mgagne: thanks for that, ill have a look
15:55:24 melwitt generally speaking, I think pack is the more desired behavior for efficient usage of compute hosts. the only reason people increased subset size, as I understand it, was to avoid the racing of parallel requests trying to claim the same nodes with the old way of claiming
15:55:46 mnaser yeah that was what we had to do for a while
15:55:47 mriedem bauzas: we call the scheduler if the server is being rebuilt with a new image
15:55:56 mriedem b/c we need to validate the new image for the host that the instance is on
15:56:10 bauzas ok, I didn't remember all the conditionals
15:56:12 melwitt yeah, if it's a first time ever image. but when I worked at yahoo we used to warm the cache for images on the compute hosts before letting users at it
15:56:25 bauzas mriedem: thanks
15:56:29 mnaser with ceph it's not even an issue because cow

Earlier   Later