Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-08
09:00:17 bauzas given it was in the filter scheduler code, it was only for that driver
09:00:26 gibi bauzas: then I' don't know why we have so much code in the scheduler manager to distinguish between the different cases
09:00:40 bauzas probably a consequence of reviews and whack-a-mole gaming
09:00:42 gibi either we got a non empty result from placement or it is a NoValidHost
09:00:56 bauzas or we could want to be more gentle
09:01:04 bauzas and in that case, it's a feature
09:01:28 gibi then my patch breaking that feature
09:01:29 bauzas but the fact is, we came from a fact in Ocata where any problem communicating to Placement was leading to a NoValidHost
09:03:40 gibi and as in Ocata placement was not mandatory we needed to easy that up
09:04:14 gibi but can we be strict in Pike?
09:09:19 bauzas gibi: no, Placement was optional in Newton but mandatory in Pike
09:09:26 bauzas graaah
09:09:35 bauzas lemme rephrase it
09:10:06 bauzas gibi: Placement = {"Newton": "optional", "Ocata": "mandatory"}
09:10:34 gibi so we can be strict in Pike
09:10:53 gibi then I'm happy to raise NoValidHost in all three cases
09:11:23 bauzas gibi: https://docs.openstack.org/releasenotes/nova/ocata.html#id7
09:11:35 bauzas "he Nova FilterScheduler driver is now able to make scheduling decisions based on the new Placement RESTful API endpoint that becomes mandatory in Ocata. "
09:12:14 bauzas and
09:12:16 bauzas " nova-scheduler process is now calling the placement API in order to get a list of valid destinations before calling the filters. That works only if all your compute nodes are fully upgraded to Ocata. If some nodes are not upgraded, the scheduler will still lookup from the DB instead which is less performant. "
09:12:49 bauzas Since the Placement service is now mandatory in Ocata, you need to deploy it and amend your compute node configuration with correct placement instructions before restarting nova-"compute or the compute node will refuse to start. "
09:12:54 bauzas anyway
09:13:21 bauzas gibi: last point for your knowledge
09:13:24 bauzas If by Newton (14.0.0), you don’t use any of the CoreFilter, RamFilter or DiskFilter, then please modify all your compute node’s configuration by amending either cpu_allocation_ratio (if you don’t use CoreFilter) or ram_allocation_ratio (if you don’t use RamFilter) or disk_allocation_ratio (if you don’t use DiskFilter) by putting a 9999.0 value for the ratio before upgrading the nova-scheduler to Ocata.
09:13:49 gibi bauzas: thanks. I'm convinced
09:13:59 bauzas the last one is error-prone
09:14:37 bauzas it was because some operators were *not* placing based on those legacy resources
09:15:22 bauzas so, in that case, in order to keep them untied with those resources, they have to fake an "infinite" allocation ratio
09:15:59 bauzas mostly FYI
09:17:51 gibi I see
09:19:01 gibi meanwhile I think I found why the server_group func tests fails in the chance scheduler removal patch
09:19:37 gibi scheduler.utils caches information about loaded filters
09:19:38 gibi https://github.com/openstack/nova/blob/master/nova/scheduler/utils.py#L397
09:19:45 bauzas gibi: that was one of my thoughts I didn't yet commented : why do you need to update test_server_group since it doesn't use chance ?
09:20:16 gibi the problem is that now we load only the ComputeFilter in other tests and that cache makes the tests interdependent
09:20:33 gibi so I have to invalidate the cache in the server_group tests
09:21:01 gibi I will push that soon
09:24:36 bauzas gibi: sounds like a fixture to me
09:25:55 gibi bauzas: like a scheduler fixture that 1) invalidate the cache 2) configure the requested filters 3) starts the scheduler service ?
09:26:28 gibi please note that the cache only affect the server group behavior
09:30:34 bauzas gibi: it affected the server group behaviour because we don't define filters to run ?
09:33:18 gibi bauzas: if the first test that hit https://github.com/openstack/nova/blob/master/nova/scheduler/utils.py#L403 has only ComputeFilter configured then a later server_group test fails as _get_group_details raise an exception based on a stale cache
09:34:18 gibi bauzas: previously every filter was configured by default, now it is on ComputeFilter
09:34:49 openstackgerrit Balazs Gibizer proposed openstack/nova master: replace chance with filter scheduler in func tests https://review.openstack.org/491529
09:35:07 gibi bauzas: ^^ now this should pass the func test on the gate as well
09:41:07 openstackgerrit Rodolfo Alonso Hernandez proposed openstack/nova master: Add datapath type information to OVS vif objects https://review.openstack.org/474892
09:43:17 bauzas gibi: I left some comments
09:43:31 bauzas gibi: I'm unclear on some needed modifications
09:43:45 gibi bauzas: looking
09:48:05 gibi bauzas: responeded
09:51:53 bauzas gibi: okay, looking
09:54:05 bauzas gibi: about the doubled allocation, I thought we were just doing that against different hosts
09:54:19 bauzas gibi: if we resize on same host, we also duplicate the allocation ?
09:54:40 bauzas :/
09:54:55 gibi bauzas: I think so https://review.openstack.org/#/c/490085/
09:55:09 bauzas gibi: could you just split your change in twice then ?
09:55:14 bauzas gibi: and test
09:56:02 gibi bauzas: do you mean one patch for the vcpu=2 and the other is the rest?
09:56:33 gibi I'm going to eat something now but then I will be back
09:57:06 bauzas gibi: well, the problem I see is that if you need to resize an instance, you absolutely now need 2 CPUs
09:57:20 bauzas even for an AIO
09:57:46 bauzas that's probably something I wasn't really concerned, but you can play with allocation ratios
09:57:57 bauzas gibi: oh had a thought
09:58:25 bauzas gibi: what if instead of modifying the fake driver resource, you would just amend the according resize test by providing a cpu allocation ?
09:58:35 bauzas allocation ratio ?
09:58:54 bauzas gibi: it should anyway default to 16.0 so I don't really see *why* we need that
10:52:11 gibi bauzas: OK, I will look into the resize test
11:04:16 cdent gibi: what new bugs have you found today
11:06:38 cdent gibi, bauzas : have you guys seen this https://review.openstack.org/#/c/489205/ is a fix for https://bugs.launchpad.net/nova/+bug/1708978 which is something we ought to make sure is in pike
11:06:39 openstack Launchpad bug 1708978 in OpenStack Compute (nova) "The traits associations are deleted incorrectly" [High,In progress] - Assigned to Alex Xu (xuhj)
11:09:05 gibi cdent: hi! no new bug today
11:09:28 cdent <- disbelief
11:09:33 gibi slow day
11:10:00 gibi still working on the removal of the change scheduler in the func test and the resize to too big flavor patches
11:10:24 gibi so I had no time to play with some custom resource + resize tests
11:10:33 gibi that will be my next fun
11:11:45 cdent I’m not having the best success trying to keep track of everything: there are lots of indvidual patch sets spread around. I need to take some time to find them all.
11:13:41 openstackgerrit Rodolfo Alonso Hernandez proposed openstack/os-vif master: Add memoize function using oslo.cache https://review.openstack.org/472773
11:13:47 gibi most of them tight to a bug report
11:14:05 gibi so if you look at the high prio bugs then your will find relevant patches
11:14:55 cdent gibi: yeah, I know, it is more in terms of being able to have them all at once for a) an overview of what’s up, b) some local testing with the pending stuff
11:15:37 cdent since they are all spread around, there’s no easy way, to, for example, answer the question of “do these fixes play well together” or “what coverage is missing”
11:17:04 gibi cdent: ahh I see
11:17:12 gibi cdent: I have no good answer for that
11:17:15 cdent :)
11:17:58 cdent I’m currently looking at coverage results when running just functional/test_servers.py to see if that raises any alarms. But because I’m looking at master I now it is missing several of the things that are in progress.
11:20:50 cdent gibi: a lot of what is missing is related to custom resource classes, so your plans for that will be useful
11:23:31 openstackgerrit Alex Xu proposed openstack/nova master: placement: ensure sharing RPs maps combinates with correct shared RP https://review.openstack.org/480379
11:23:32 openstackgerrit Balazs Gibizer proposed openstack/nova master: replace chance with filter scheduler in func tests https://review.openstack.org/491529
11:23:52 alex_xu cdent: ^ remove the 'root', instead to use 'sharing' and 'shared'
11:24:02 gibi cdent: now I just have to find the time to write them :)
11:24:08 cdent thank you alex_xu
11:24:27 gibi bauzas: I removed the vcpu=2 from the SmallFakeDriver to see what fails
11:24:46 gibi bauzas: I think we have a problem with the default 16.0 allocation ration. I don't see that it is applied at all
11:24:58 alex_xu cdent: hope that works :)
11:26:26 gibi bauzas: here is an example test failure http://paste.openstack.org/show/617764/
11:26:26 cdent gibi: do you get reasonable results from placement, but then the fake driver refuses then? If so, it’s probably a bug in the driver itself. If you’re not getting results from placement then is the inventory being set properly?
11:27:39 gibi cdent: L66 worries me http://paste.openstack.org/show/617764/
11:28:15 cdent gibi: it’s max_unit,
11:28:21 cdent that’s the problem

Earlier   Later