| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-08 | |||
| 08:34:29 | bauzas | tbc, the chance one | |
| 08:34:48 | bauzas | did Gerrit was upgraded ? | |
| 08:35:00 | bauzas | I just provided an URL in a comment and it broke weirdly | |
| 08:36:10 | gibi | bauzas: that change still fail on the gate with the server_group tests but I cannot repoduce it locally and I even removed every server_group related change for them patch | |
| 08:36:22 | bauzas | gibi: ok | |
| 08:37:04 | gibi | bauzas: so any comment is really appreciated | |
| 08:40:37 | gibi | bauzas: also I'm curious about your oppinion about https://review.openstack.org/#/c/491491/4/nova/tests/unit/scheduler/test_scheduler.py@152 | |
| 08:41:40 | bauzas | gibi: you mean about legacy filters being still in use ? | |
| 08:42:14 | gibi | bauzas: about raising NoValidHost if Placement API is not available or not upgraded | |
| 08:43:23 | gibi | bauzas: raising NoValidHost if Placemnet returns {} is clearly a good thing but I'm a bit uneasy about the two other cases | |
| 08:49:04 | bauzas | gibi: sorry, was just refueling my stomach | |
| 08:49:23 | bauzas | by coffee (10:40am here, what people are thinking ?) | |
| 08:50:09 | bauzas | gibi: about the NoValidHost exception ? Well, correct me if I'm wrong but we already do that :) | |
| 08:50:22 | bauzas | gibi: we just return an empty list that will eventually raise that error :) | |
| 08:53:07 | gibi | bauzas: if Placement is not available then Scheduler falls back to normal filtering and if CoreFilter is not enabled it can select a host with not enough vcpu | |
| 08:54:24 | gibi | bauzas: now this can happen in three cases, Placement cannot be accessed, Placement is not upgraded, or Placement returned {}. | |
| 08:54:44 | gibi | bauzas: in the third one I'm happy to raise NoValidHost | |
| 08:54:48 | bauzas | gibi: wait, are you sure ? | |
| 08:55:30 | bauzas | gibi: when I wrote the BP about scheduler calling placement in Ocata, I made it clear that if Placement wasn't there, we should return an empty list | |
| 08:55:31 | gibi | bauzas: based on this https://github.com/openstack/nova/blob/master/nova/scheduler/manager.py#L133 | |
| 08:55:44 | bauzas | then it's a regression | |
| 08:55:56 | bauzas | in Pike | |
| 08:56:25 | gibi | so it is OK for you to raise NoValidHost in all the 3 cases. I can accept this. :) | |
| 08:56:27 | bauzas | Placement isn't optional | |
| 08:56:34 | bauzas | if the driver supports it | |
| 08:56:50 | bauzas | at least that's what we had in Ocata | |
| 08:56:57 | bauzas | gibi: lemme show you the ocata code | |
| 08:57:12 | gibi | bauzas: but then this whole bunch of code is unnecessary https://github.com/openstack/nova/blob/master/nova/scheduler/manager.py#L124-L144 | |
| 08:58:27 | bauzas | gibi: https://github.com/openstack/nova/blob/stable/ocata/nova/scheduler/filter_scheduler.py#L184-L189 | |
| 08:59:31 | gibi | bauzas: OK, that sounds convincing | |
| 09:00:00 | bauzas | that's at least what we had in Ocata | |
| 09:00:17 | bauzas | given it was in the filter scheduler code, it was only for that driver | |
| 09:00:26 | gibi | bauzas: then I' don't know why we have so much code in the scheduler manager to distinguish between the different cases | |
| 09:00:40 | bauzas | probably a consequence of reviews and whack-a-mole gaming | |
| 09:00:42 | gibi | either we got a non empty result from placement or it is a NoValidHost | |
| 09:00:56 | bauzas | or we could want to be more gentle | |
| 09:01:04 | bauzas | and in that case, it's a feature | |
| 09:01:28 | gibi | then my patch breaking that feature | |
| 09:01:29 | bauzas | but the fact is, we came from a fact in Ocata where any problem communicating to Placement was leading to a NoValidHost | |
| 09:03:40 | gibi | and as in Ocata placement was not mandatory we needed to easy that up | |
| 09:04:14 | gibi | but can we be strict in Pike? | |
| 09:09:19 | bauzas | gibi: no, Placement was optional in Newton but mandatory in Pike | |
| 09:09:26 | bauzas | graaah | |
| 09:09:35 | bauzas | lemme rephrase it | |
| 09:10:06 | bauzas | gibi: Placement = {"Newton": "optional", "Ocata": "mandatory"} | |
| 09:10:34 | gibi | so we can be strict in Pike | |
| 09:10:53 | gibi | then I'm happy to raise NoValidHost in all three cases | |
| 09:11:23 | bauzas | gibi: https://docs.openstack.org/releasenotes/nova/ocata.html#id7 | |
| 09:11:35 | bauzas | "he Nova FilterScheduler driver is now able to make scheduling decisions based on the new Placement RESTful API endpoint that becomes mandatory in Ocata. " | |
| 09:12:14 | bauzas | and | |
| 09:12:16 | bauzas | " nova-scheduler process is now calling the placement API in order to get a list of valid destinations before calling the filters. That works only if all your compute nodes are fully upgraded to Ocata. If some nodes are not upgraded, the scheduler will still lookup from the DB instead which is less performant. " | |
| 09:12:49 | bauzas | Since the Placement service is now mandatory in Ocata, you need to deploy it and amend your compute node configuration with correct placement instructions before restarting nova-"compute or the compute node will refuse to start. " | |
| 09:12:54 | bauzas | anyway | |
| 09:13:21 | bauzas | gibi: last point for your knowledge | |
| 09:13:24 | bauzas | If by Newton (14.0.0), you don’t use any of the CoreFilter, RamFilter or DiskFilter, then please modify all your compute node’s configuration by amending either cpu_allocation_ratio (if you don’t use CoreFilter) or ram_allocation_ratio (if you don’t use RamFilter) or disk_allocation_ratio (if you don’t use DiskFilter) by putting a 9999.0 value for the ratio before upgrading the nova-scheduler to Ocata. | |
| 09:13:49 | gibi | bauzas: thanks. I'm convinced | |
| 09:13:59 | bauzas | the last one is error-prone | |
| 09:14:37 | bauzas | it was because some operators were *not* placing based on those legacy resources | |
| 09:15:22 | bauzas | so, in that case, in order to keep them untied with those resources, they have to fake an "infinite" allocation ratio | |
| 09:15:59 | bauzas | mostly FYI | |
| 09:17:51 | gibi | I see | |
| 09:19:01 | gibi | meanwhile I think I found why the server_group func tests fails in the chance scheduler removal patch | |
| 09:19:37 | gibi | scheduler.utils caches information about loaded filters | |
| 09:19:38 | gibi | https://github.com/openstack/nova/blob/master/nova/scheduler/utils.py#L397 | |
| 09:19:45 | bauzas | gibi: that was one of my thoughts I didn't yet commented : why do you need to update test_server_group since it doesn't use chance ? | |
| 09:20:16 | gibi | the problem is that now we load only the ComputeFilter in other tests and that cache makes the tests interdependent | |
| 09:20:33 | gibi | so I have to invalidate the cache in the server_group tests | |
| 09:21:01 | gibi | I will push that soon | |
| 09:24:36 | bauzas | gibi: sounds like a fixture to me | |
| 09:25:55 | gibi | bauzas: like a scheduler fixture that 1) invalidate the cache 2) configure the requested filters 3) starts the scheduler service ? | |
| 09:26:28 | gibi | please note that the cache only affect the server group behavior | |
| 09:30:34 | bauzas | gibi: it affected the server group behaviour because we don't define filters to run ? | |
| 09:33:18 | gibi | bauzas: if the first test that hit https://github.com/openstack/nova/blob/master/nova/scheduler/utils.py#L403 has only ComputeFilter configured then a later server_group test fails as _get_group_details raise an exception based on a stale cache | |
| 09:34:18 | gibi | bauzas: previously every filter was configured by default, now it is on ComputeFilter | |
| 09:34:49 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: replace chance with filter scheduler in func tests https://review.openstack.org/491529 | |
| 09:35:07 | gibi | bauzas: ^^ now this should pass the func test on the gate as well | |
| 09:41:07 | openstackgerrit | Rodolfo Alonso Hernandez proposed openstack/nova master: Add datapath type information to OVS vif objects https://review.openstack.org/474892 | |
| 09:43:17 | bauzas | gibi: I left some comments | |
| 09:43:31 | bauzas | gibi: I'm unclear on some needed modifications | |
| 09:43:45 | gibi | bauzas: looking | |
| 09:48:05 | gibi | bauzas: responeded | |
| 09:51:53 | bauzas | gibi: okay, looking | |
| 09:54:05 | bauzas | gibi: about the doubled allocation, I thought we were just doing that against different hosts | |
| 09:54:19 | bauzas | gibi: if we resize on same host, we also duplicate the allocation ? | |
| 09:54:40 | bauzas | :/ | |
| 09:54:55 | gibi | bauzas: I think so https://review.openstack.org/#/c/490085/ | |
| 09:55:09 | bauzas | gibi: could you just split your change in twice then ? | |
| 09:55:14 | bauzas | gibi: and test | |
| 09:56:02 | gibi | bauzas: do you mean one patch for the vcpu=2 and the other is the rest? | |
| 09:56:33 | gibi | I'm going to eat something now but then I will be back | |
| 09:57:06 | bauzas | gibi: well, the problem I see is that if you need to resize an instance, you absolutely now need 2 CPUs | |
| 09:57:20 | bauzas | even for an AIO | |
| 09:57:46 | bauzas | that's probably something I wasn't really concerned, but you can play with allocation ratios | |
| 09:57:57 | bauzas | gibi: oh had a thought | |
| 09:58:25 | bauzas | gibi: what if instead of modifying the fake driver resource, you would just amend the according resize test by providing a cpu allocation ? | |
| 09:58:35 | bauzas | allocation ratio ? | |
| 09:58:54 | bauzas | gibi: it should anyway default to 16.0 so I don't really see *why* we need that | |
| 10:52:11 | gibi | bauzas: OK, I will look into the resize test | |
| 11:04:16 | cdent | gibi: what new bugs have you found today | |
| 11:06:38 | cdent | gibi, bauzas : have you guys seen this https://review.openstack.org/#/c/489205/ is a fix for https://bugs.launchpad.net/nova/+bug/1708978 which is something we ought to make sure is in pike | |