| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-08 | |||
| 09:04:14 | gibi | but can we be strict in Pike? | |
| 09:09:19 | bauzas | gibi: no, Placement was optional in Newton but mandatory in Pike | |
| 09:09:26 | bauzas | graaah | |
| 09:09:35 | bauzas | lemme rephrase it | |
| 09:10:06 | bauzas | gibi: Placement = {"Newton": "optional", "Ocata": "mandatory"} | |
| 09:10:34 | gibi | so we can be strict in Pike | |
| 09:10:53 | gibi | then I'm happy to raise NoValidHost in all three cases | |
| 09:11:23 | bauzas | gibi: https://docs.openstack.org/releasenotes/nova/ocata.html#id7 | |
| 09:11:35 | bauzas | "he Nova FilterScheduler driver is now able to make scheduling decisions based on the new Placement RESTful API endpoint that becomes mandatory in Ocata. " | |
| 09:12:14 | bauzas | and | |
| 09:12:16 | bauzas | " nova-scheduler process is now calling the placement API in order to get a list of valid destinations before calling the filters. That works only if all your compute nodes are fully upgraded to Ocata. If some nodes are not upgraded, the scheduler will still lookup from the DB instead which is less performant. " | |
| 09:12:49 | bauzas | Since the Placement service is now mandatory in Ocata, you need to deploy it and amend your compute node configuration with correct placement instructions before restarting nova-"compute or the compute node will refuse to start. " | |
| 09:12:54 | bauzas | anyway | |
| 09:13:21 | bauzas | gibi: last point for your knowledge | |
| 09:13:24 | bauzas | If by Newton (14.0.0), you don’t use any of the CoreFilter, RamFilter or DiskFilter, then please modify all your compute node’s configuration by amending either cpu_allocation_ratio (if you don’t use CoreFilter) or ram_allocation_ratio (if you don’t use RamFilter) or disk_allocation_ratio (if you don’t use DiskFilter) by putting a 9999.0 value for the ratio before upgrading the nova-scheduler to Ocata. | |
| 09:13:49 | gibi | bauzas: thanks. I'm convinced | |
| 09:13:59 | bauzas | the last one is error-prone | |
| 09:14:37 | bauzas | it was because some operators were *not* placing based on those legacy resources | |
| 09:15:22 | bauzas | so, in that case, in order to keep them untied with those resources, they have to fake an "infinite" allocation ratio | |
| 09:15:59 | bauzas | mostly FYI | |
| 09:17:51 | gibi | I see | |
| 09:19:01 | gibi | meanwhile I think I found why the server_group func tests fails in the chance scheduler removal patch | |
| 09:19:37 | gibi | scheduler.utils caches information about loaded filters | |
| 09:19:38 | gibi | https://github.com/openstack/nova/blob/master/nova/scheduler/utils.py#L397 | |
| 09:19:45 | bauzas | gibi: that was one of my thoughts I didn't yet commented : why do you need to update test_server_group since it doesn't use chance ? | |
| 09:20:16 | gibi | the problem is that now we load only the ComputeFilter in other tests and that cache makes the tests interdependent | |
| 09:20:33 | gibi | so I have to invalidate the cache in the server_group tests | |
| 09:21:01 | gibi | I will push that soon | |
| 09:24:36 | bauzas | gibi: sounds like a fixture to me | |
| 09:25:55 | gibi | bauzas: like a scheduler fixture that 1) invalidate the cache 2) configure the requested filters 3) starts the scheduler service ? | |
| 09:26:28 | gibi | please note that the cache only affect the server group behavior | |
| 09:30:34 | bauzas | gibi: it affected the server group behaviour because we don't define filters to run ? | |
| 09:33:18 | gibi | bauzas: if the first test that hit https://github.com/openstack/nova/blob/master/nova/scheduler/utils.py#L403 has only ComputeFilter configured then a later server_group test fails as _get_group_details raise an exception based on a stale cache | |
| 09:34:18 | gibi | bauzas: previously every filter was configured by default, now it is on ComputeFilter | |
| 09:34:49 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: replace chance with filter scheduler in func tests https://review.openstack.org/491529 | |
| 09:35:07 | gibi | bauzas: ^^ now this should pass the func test on the gate as well | |
| 09:41:07 | openstackgerrit | Rodolfo Alonso Hernandez proposed openstack/nova master: Add datapath type information to OVS vif objects https://review.openstack.org/474892 | |
| 09:43:17 | bauzas | gibi: I left some comments | |
| 09:43:31 | bauzas | gibi: I'm unclear on some needed modifications | |
| 09:43:45 | gibi | bauzas: looking | |
| 09:48:05 | gibi | bauzas: responeded | |
| 09:51:53 | bauzas | gibi: okay, looking | |
| 09:54:05 | bauzas | gibi: about the doubled allocation, I thought we were just doing that against different hosts | |
| 09:54:19 | bauzas | gibi: if we resize on same host, we also duplicate the allocation ? | |
| 09:54:40 | bauzas | :/ | |
| 09:54:55 | gibi | bauzas: I think so https://review.openstack.org/#/c/490085/ | |
| 09:55:09 | bauzas | gibi: could you just split your change in twice then ? | |
| 09:55:14 | bauzas | gibi: and test | |
| 09:56:02 | gibi | bauzas: do you mean one patch for the vcpu=2 and the other is the rest? | |
| 09:56:33 | gibi | I'm going to eat something now but then I will be back | |
| 09:57:06 | bauzas | gibi: well, the problem I see is that if you need to resize an instance, you absolutely now need 2 CPUs | |
| 09:57:20 | bauzas | even for an AIO | |
| 09:57:46 | bauzas | that's probably something I wasn't really concerned, but you can play with allocation ratios | |
| 09:57:57 | bauzas | gibi: oh had a thought | |
| 09:58:25 | bauzas | gibi: what if instead of modifying the fake driver resource, you would just amend the according resize test by providing a cpu allocation ? | |
| 09:58:35 | bauzas | allocation ratio ? | |
| 09:58:54 | bauzas | gibi: it should anyway default to 16.0 so I don't really see *why* we need that | |
| 10:52:11 | gibi | bauzas: OK, I will look into the resize test | |
| 11:04:16 | cdent | gibi: what new bugs have you found today | |
| 11:06:38 | cdent | gibi, bauzas : have you guys seen this https://review.openstack.org/#/c/489205/ is a fix for https://bugs.launchpad.net/nova/+bug/1708978 which is something we ought to make sure is in pike | |
| 11:06:39 | openstack | Launchpad bug 1708978 in OpenStack Compute (nova) "The traits associations are deleted incorrectly" [High,In progress] - Assigned to Alex Xu (xuhj) | |
| 11:09:05 | gibi | cdent: hi! no new bug today | |
| 11:09:28 | cdent | <- disbelief | |
| 11:09:33 | gibi | slow day | |
| 11:10:00 | gibi | still working on the removal of the change scheduler in the func test and the resize to too big flavor patches | |
| 11:10:24 | gibi | so I had no time to play with some custom resource + resize tests | |
| 11:10:33 | gibi | that will be my next fun | |
| 11:11:45 | cdent | I’m not having the best success trying to keep track of everything: there are lots of indvidual patch sets spread around. I need to take some time to find them all. | |
| 11:13:41 | openstackgerrit | Rodolfo Alonso Hernandez proposed openstack/os-vif master: Add memoize function using oslo.cache https://review.openstack.org/472773 | |
| 11:13:47 | gibi | most of them tight to a bug report | |
| 11:14:05 | gibi | so if you look at the high prio bugs then your will find relevant patches | |
| 11:14:55 | cdent | gibi: yeah, I know, it is more in terms of being able to have them all at once for a) an overview of what’s up, b) some local testing with the pending stuff | |
| 11:15:37 | cdent | since they are all spread around, there’s no easy way, to, for example, answer the question of “do these fixes play well together” or “what coverage is missing” | |
| 11:17:04 | gibi | cdent: ahh I see | |
| 11:17:12 | gibi | cdent: I have no good answer for that | |
| 11:17:15 | cdent | :) | |
| 11:17:58 | cdent | I’m currently looking at coverage results when running just functional/test_servers.py to see if that raises any alarms. But because I’m looking at master I now it is missing several of the things that are in progress. | |
| 11:20:50 | cdent | gibi: a lot of what is missing is related to custom resource classes, so your plans for that will be useful | |
| 11:23:31 | openstackgerrit | Alex Xu proposed openstack/nova master: placement: ensure sharing RPs maps combinates with correct shared RP https://review.openstack.org/480379 | |
| 11:23:32 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: replace chance with filter scheduler in func tests https://review.openstack.org/491529 | |
| 11:23:52 | alex_xu | cdent: ^ remove the 'root', instead to use 'sharing' and 'shared' | |
| 11:24:02 | gibi | cdent: now I just have to find the time to write them :) | |
| 11:24:08 | cdent | thank you alex_xu | |
| 11:24:27 | gibi | bauzas: I removed the vcpu=2 from the SmallFakeDriver to see what fails | |
| 11:24:46 | gibi | bauzas: I think we have a problem with the default 16.0 allocation ration. I don't see that it is applied at all | |
| 11:24:58 | alex_xu | cdent: hope that works :) | |
| 11:26:26 | gibi | bauzas: here is an example test failure http://paste.openstack.org/show/617764/ | |
| 11:26:26 | cdent | gibi: do you get reasonable results from placement, but then the fake driver refuses then? If so, it’s probably a bug in the driver itself. If you’re not getting results from placement then is the inventory being set properly? | |
| 11:27:39 | gibi | cdent: L66 worries me http://paste.openstack.org/show/617764/ | |
| 11:28:15 | cdent | gibi: it’s max_unit, | |
| 11:28:21 | cdent | that’s the problem | |
| 11:28:31 | cdent | and is an actual problem: | |
| 11:29:10 | cdent | we set max_unit to be the number of real cpus because we don’t think any single consumer should occupy more than the number of cpus, whatever allocation ratio says | |
| 11:29:13 | cdent | which makes sense | |
| 11:29:31 | cdent | but when we create a doubling allocation for the resize to same host, we’re breaking that | |
| 11:30:13 | gibi | bauzas: ^^ | |
| 11:30:17 | cdent | so as the code is currently designed, if using doubling, you can never resize to same host a guest with #vcpus == #pcpus | |
| 11:30:29 | cdent | the quick fix for the tests is to use a bigger driverr | |
| 11:30:44 | gibi | yes, that was what I proposed the bauzas suggested allocation ratio | |
| 11:30:54 | gibi | but then I guess we cannot use the allocation ratio trick here | |