| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-08 | |||
| 11:09:33 | gibi | slow day | |
| 11:10:00 | gibi | still working on the removal of the change scheduler in the func test and the resize to too big flavor patches | |
| 11:10:24 | gibi | so I had no time to play with some custom resource + resize tests | |
| 11:10:33 | gibi | that will be my next fun | |
| 11:11:45 | cdent | I’m not having the best success trying to keep track of everything: there are lots of indvidual patch sets spread around. I need to take some time to find them all. | |
| 11:13:41 | openstackgerrit | Rodolfo Alonso Hernandez proposed openstack/os-vif master: Add memoize function using oslo.cache https://review.openstack.org/472773 | |
| 11:13:47 | gibi | most of them tight to a bug report | |
| 11:14:05 | gibi | so if you look at the high prio bugs then your will find relevant patches | |
| 11:14:55 | cdent | gibi: yeah, I know, it is more in terms of being able to have them all at once for a) an overview of what’s up, b) some local testing with the pending stuff | |
| 11:15:37 | cdent | since they are all spread around, there’s no easy way, to, for example, answer the question of “do these fixes play well together” or “what coverage is missing” | |
| 11:17:04 | gibi | cdent: ahh I see | |
| 11:17:12 | gibi | cdent: I have no good answer for that | |
| 11:17:15 | cdent | :) | |
| 11:17:58 | cdent | I’m currently looking at coverage results when running just functional/test_servers.py to see if that raises any alarms. But because I’m looking at master I now it is missing several of the things that are in progress. | |
| 11:20:50 | cdent | gibi: a lot of what is missing is related to custom resource classes, so your plans for that will be useful | |
| 11:23:31 | openstackgerrit | Alex Xu proposed openstack/nova master: placement: ensure sharing RPs maps combinates with correct shared RP https://review.openstack.org/480379 | |
| 11:23:32 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: replace chance with filter scheduler in func tests https://review.openstack.org/491529 | |
| 11:23:52 | alex_xu | cdent: ^ remove the 'root', instead to use 'sharing' and 'shared' | |
| 11:24:02 | gibi | cdent: now I just have to find the time to write them :) | |
| 11:24:08 | cdent | thank you alex_xu | |
| 11:24:27 | gibi | bauzas: I removed the vcpu=2 from the SmallFakeDriver to see what fails | |
| 11:24:46 | gibi | bauzas: I think we have a problem with the default 16.0 allocation ration. I don't see that it is applied at all | |
| 11:24:58 | alex_xu | cdent: hope that works :) | |
| 11:26:26 | cdent | gibi: do you get reasonable results from placement, but then the fake driver refuses then? If so, it’s probably a bug in the driver itself. If you’re not getting results from placement then is the inventory being set properly? | |
| 11:26:26 | gibi | bauzas: here is an example test failure http://paste.openstack.org/show/617764/ | |
| 11:27:39 | gibi | cdent: L66 worries me http://paste.openstack.org/show/617764/ | |
| 11:28:15 | cdent | gibi: it’s max_unit, | |
| 11:28:21 | cdent | that’s the problem | |
| 11:28:31 | cdent | and is an actual problem: | |
| 11:29:10 | cdent | we set max_unit to be the number of real cpus because we don’t think any single consumer should occupy more than the number of cpus, whatever allocation ratio says | |
| 11:29:13 | cdent | which makes sense | |
| 11:29:31 | cdent | but when we create a doubling allocation for the resize to same host, we’re breaking that | |
| 11:30:13 | gibi | bauzas: ^^ | |
| 11:30:17 | cdent | so as the code is currently designed, if using doubling, you can never resize to same host a guest with #vcpus == #pcpus | |
| 11:30:29 | cdent | the quick fix for the tests is to use a bigger driverr | |
| 11:30:44 | gibi | yes, that was what I proposed the bauzas suggested allocation ratio | |
| 11:30:54 | gibi | but then I guess we cannot use the allocation ratio trick here | |
| 11:31:02 | cdent | but we may need to consider that the doubling concept is problematic on small hosts… | |
| 11:31:14 | cdent | yeah, allocation ratio doesn’t do anything for max_unit | |
| 11:32:34 | gibi | OK I wait for bauzas to agree then I will push back the previos patch set of https://review.openstack.org/#/c/491529/ that has the SmallFakeDriver modification to 2 vcpus | |
| 11:33:07 | cdent | alex_xu: it isn’t quite right. I’m trying to come up with a suitable alternative. I’ll push something up if that’s okay with you? | |
| 11:33:53 | alex_xu | cdent: yea, appreciate the help, looks like I understand 'shared' and 'sharing' incorrectly | |
| 11:33:58 | openstackgerrit | Sean Dague proposed openstack/nova master: Fix all >= 2 hit 404s https://review.openstack.org/491761 | |
| 11:34:27 | cdent | gibi: seems reasonable. I wonder who has an answer to the question about the legitimacy of resizing a large guest on a small host. I don’t really know. I wouldn’t want to do it, but I’m not paying for hardware... | |
| 11:35:44 | gibi | resize on same host feels like a feature to support testing resize on single node devstack | |
| 11:39:42 | openstackgerrit | Sean Dague proposed openstack/nova master: Create reference subpage https://review.openstack.org/490994 | |
| 11:51:45 | openstackgerrit | Chris Dent proposed openstack/nova master: placement: ensure RP maps to those RPs that share with it https://review.openstack.org/480379 | |
| 11:54:12 | cdent | alex_xu: ^ that may be a bit better. I ended up finding it difficult to make clear. | |
| 11:55:33 | alex_xu | cdent: thanks | |
| 11:58:22 | openstackgerrit | OpenStack Proposal Bot proposed openstack/nova master: Imported Translations from Zanata https://review.openstack.org/477091 | |
| 12:01:24 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Raise NoValidHost if no allocation candidates https://review.openstack.org/491491 | |
| 12:02:31 | gibi | bauzas: fixed your comment about the test ^^ | |
| 12:12:33 | bauzas | cdent: looking at alex_xu change, thanks for it | |
| 12:13:32 | bauzas | gibi: cdent: looking at the above discussion, could you please summarize the problem with ratios ? | |
| 12:14:23 | gibi | bauzas: we set max_unit of the vcpu resource based on the number of pcpus therefore we cannot ask for more than pcpu amount of vcpu in a single allocation | |
| 12:14:39 | alex_xu | bauzas: thanks for the review | |
| 12:14:48 | bauzas | gibi: mmm, sadly then | |
| 12:21:17 | gibi | bauzas: so are you OK with the SmallFakeDriver vcpu=2 change? | |
| 12:21:24 | bauzas | gibi: looking | |
| 12:23:39 | bauzas | gibi: just reviewing alex_xu and then it's your turn :p | |
| 12:25:57 | gibi | bauzas: OK, thanks | |
| 12:28:53 | openstackgerrit | Alexandra Settle proposed openstack/nova master: Create reference subpage https://review.openstack.org/490994 | |
| 12:29:21 | bauzas | cdent: gibi: so, question about max_unit | |
| 12:29:47 | bauzas | cdent: gibi: when we have a flavor asking for 10 CPUs, do you think it should be 10 pCPUs or 10 vCPUs ? | |
| 12:30:01 | bauzas | cdent: gibi: MHO is that's for virtual ones | |
| 12:30:02 | asettle | Ah shite, sdague and stephenfin - I got ahead of myself and updated the wrong nova patch (didn't tab to the right PR). I just updated this patch: https://review.openstack.org/#/c/490994/ with the change from Administrators to Adminstration Guide | |
| 12:30:04 | asettle | I can revert if you'd like? | |
| 12:30:14 | cdent | bauzas: it _is_ for vcpus | |
| 12:30:18 | asettle | Otherwise, I was just going to make the same change on stephenfin 's patch: https://review.openstack.org/#/c/490952/ | |
| 12:30:21 | bauzas | cdent: I know | |
| 12:30:30 | cdent | and it should be | |
| 12:30:41 | bauzas | cdent: gibi: but the point is then, why max_unit should be about *physical* resources ? | |
| 12:31:56 | cdent | max_unit, in general, is abstract. where we make concrete assertions about its meaning is when we set inventory from the compute manager. In there we make the assertion that it is bad for a service to host a _single_ vm that has more vcpus than there are pcpus | |
| 12:32:34 | cdent | when we discussed this before, several people mentioned that doing so would cause pathalogical context switching | |
| 12:32:44 | bauzas | cdent: tbc, flavor * ratio should be <= max_unit | |
| 12:33:16 | cdent | ratio has _nothing_ to do with max unit. ratio is a measurement for all all guests, max_unit is for one | |
| 12:33:16 | bauzas | not sure I understand the problem about max_unit being related to physical resources :( | |
| 12:33:37 | bauzas | cdent: sure, I know | |
| 12:34:09 | cdent | bauzas: okay, so I dont understand what you’re asking or suggesting then? | |
| 12:34:16 | sdague | asettle: don't work, I've got a patch sitting on top of it that I'm about to push and it should zero it back | |
| 12:34:31 | asettle | sdague: dont work? | |
| 12:34:37 | cdent | asettle: woot. you never need to work again! | |
| 12:34:44 | asettle | cdent: I have seriously been waiting for this day for ages. | |
| 12:34:49 | bauzas | cdent: the problem is that max_unit is blocking us for https://review.openstack.org/#/c/491529/4/nova/virt/fake.py@599 | |
| 12:34:57 | sdague | asettle: don't worry | |
| 12:35:14 | cdent | bauzas: yes, in a way we want and expect it to. so either we change the fake driver or use a different one. | |
| 12:35:14 | asettle | sdague: hahaha okay :) | |
| 12:35:30 | asettle | I will go make my two second patch on stephenfin 's then | |
| 12:35:52 | bauzas | cdent: I'm unclear, why should we need to modify the resources if we have a ratio around 16.0 ? | |
| 12:36:08 | bauzas | cdent: because of max_unit, because max_unit == pCPUs, right? | |
| 12:36:30 | bauzas | that ^ looks weird to me, and changing the behaviour we had for 5 years | |
| 12:36:37 | cdent | because we are creating an allocation, when we do the doubling, for a single consumer that is greater than max_unit | |
| 12:36:49 | cdent | the change in behavior is the doubling | |
| 12:36:53 | cdent | that’s what’s new | |
| 12:37:01 | cdent | the max_unit behavior has been around for a while | |
| 12:37:13 | bauzas | imagine a world where we would double the allocation, but have the previous behaviour | |
| 12:37:23 | cdent | if we were able to do the doubling with two different consumer uuids, we wouldnt have this problem | |
| 12:37:45 | bauzas | say I have a single pCPU with 16.0 ratio, I can still ask for 2 flavors of 1vCPU, right? | |
| 12:37:53 | cdent | bauzas: yes | |
| 12:38:05 | bauzas | so, why can't we do that now ? | |