| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-09 | |||
| 15:42:52 | mnaser | numastat -m reports 4096 free hugepages in each node (for a total of 8192) | |
| 15:45:23 | cfriesen | too bad sfinucan is on holidays. :) | |
| 15:45:59 | mnaser | the used cores are: 2-15,18-31 | |
| 15:46:12 | mnaser | which leaves 16/17 free which are on two different sockets | |
| 15:46:43 | mnaser | so technically it should schedule on 16,17 which will be able to get 4096MB from each numa socket | |
| 15:47:53 | cfriesen | mnaser: which version of the code is this? | |
| 15:47:58 | mnaser | newton | |
| 15:48:09 | mnaser | python-nova-14.0.7-1.el7.noarch | |
| 15:48:10 | mnaser | more specifically | |
| 15:49:47 | mnaser | sahid i looked at your change however it seems siblings_set being empty is a possiblity :( | |
| 15:50:03 | cfriesen | dansmith: for some reason I thought that pCPUs without a sibling would still be in sibling_set, just as a single-item set | |
| 15:50:46 | efried | sdague https://github.com/openstack/nova/blob/master/nova/cmd/status.py#L198 Am I missing something, or is this untrue? | |
| 15:51:11 | hongbin | oomichi: hi, my team is working on deciding the default api version, a team member asked me to check with you since you did the api version work in nova, i mainly wanted to know how nova maintain the default api version (pick the latest version or a stable version? bump the default version everytime a new version is introduced? etc.) Are you the right person to ask? | |
| 15:51:28 | mnaser | does host_cell get assigned by the scheduler? | |
| 15:52:44 | cfriesen | mnaser: danpb: dansmith: this looks sort of related, but the fix should be in newton: https://bugs.launchpad.net/nova/+bug/1578155 | |
| 15:52:45 | openstack | Launchpad bug 1578155 in OpenStack Compute (nova) newton "'hw:cpu_thread_policy=prefer' misbehaviour" [Medium,Fix committed] - Assigned to Stephen Finucane (stephenfinucane) | |
| 15:52:57 | dansmith | I really don't know anything about this stuff | |
| 15:54:53 | jaypipes | mriedem, dansmith: I will be afk for 2 hours, just FYI. | |
| 15:54:55 | cfriesen | mnaser: so are 16/17 the siblings of 0/1 which are not included in vcpu_pin_set? | |
| 15:55:08 | openstackgerrit | Matt Riedemann proposed openstack/nova master: [placement] Add api-ref for usages https://review.openstack.org/480563 | |
| 15:55:11 | artom | About the fix, at any rate. | |
| 15:55:13 | mnaser | cfriesen nope, 0/1 are the ones which are not included | |
| 15:55:26 | mnaser | so our vcpu_pin_set is 2-31 | |
| 15:55:30 | mnaser | (32 core machine) | |
| 15:55:33 | cfriesen | mnaser: yes, but which are the HT siblings of 0/1? | |
| 15:56:15 | danpb | hmm, yes, it seems numa node 0 has even numbered cpus, and node 1 has odd numbered cpus | |
| 15:56:19 | mnaser | cfriesen sorry, not sure i follow, but here's lscpu output if that helps? http://paste.openstack.org/show/617955/ | |
| 15:56:34 | danpb | so excluding 0-1, kills 1 cpu each host numa node, instead of 2 cpus | |
| 15:56:35 | cfriesen | mnaser: based on the pattern above, the HT siblings of 0/1 would be 16/17....I'm wondering if that's confusing the logic in hardware.py | |
| 15:56:42 | mnaser | but i think that they are | |
| 15:57:19 | cfriesen | mnaswer: running "virsh capabilities" will show sibling info | |
| 15:57:26 | danpb | so that in turn kills a entire sibling pair from each numa node | |
| 15:57:37 | danpb | effectively making 4 cpus unavailable | |
| 15:57:53 | cfriesen | danpb: theoretically it shouldn't kill the sibling pair though...we should have two pCPUs with no siblings | |
| 15:58:03 | mnaser | here is output of virsh caps. http://paste.openstack.org/show/617956/ | |
| 15:58:06 | cfriesen | but maybe the logic can't handle a mix of sibs and no sibs | |
| 15:58:15 | mriedem | someone want to approve this? https://review.openstack.org/#/c/491855/ | |
| 15:58:18 | danpb | yeah that wouldn't be surprising | |
| 15:58:46 | mnaser | so am i better off reserving 2 siblings and maybe that might get rid of the 'confusion' ? | |
| 15:58:53 | cfriesen | mnaswer: seems likely | |
| 15:58:55 | mnaser | so instead of 0,1 .. do 0,16 instead? | |
| 15:59:02 | danpb | yeah | |
| 15:59:53 | mnaser | let me test that theory out | |
| 16:00:03 | openstackgerrit | Ed Leafe proposed openstack/nova master: Handle addition of new nodes/instances in ironic flavor migration https://review.openstack.org/487954 | |
| 16:00:12 | edleafe | dansmith: dtantsur: ^^ updated to remove the node RC change handling | |
| 16:00:37 | dtantsur | awesome! I'm just out of a meeting, will do the ironic part now | |
| 16:00:50 | mnaser | so my vcpu_pin_set should technically be 1-15,17-31 in that case, correct? | |
| 16:00:53 | mriedem | gibi: wasn't that also a problem in ocata then? CoreFilter isn't enabled by default | |
| 16:00:53 | mnaser | leaving 0,16 for the OS | |
| 16:00:59 | cfriesen | mnaswer: either way, please raise a bug for this issue, it'd be nice to handle it properly | |
| 16:01:02 | danpb | mnaser: yeah | |
| 16:01:11 | mnaser | cfriesen agreed | |
| 16:01:15 | mnaser | ok, lets try this out | |
| 16:01:35 | mnaser | ill reboot with the new updates isolcpus= values so itll take me a few minutes and report | |
| 16:02:31 | dansmith | mriedem: melwitt: I'm going to disappear suddenly somewhere in the hour of the cells meeting.. assume we were on track to punt anyway, but.. that okay? | |
| 16:03:57 | dansmith | jaypipes: so what's the deal with that last patch? do you want me to update it and remove the extraneous continue? or log something there or what? | |
| 16:04:15 | jaypipes | please, yes, go for it. | |
| 16:04:32 | dansmith | ...which? | |
| 16:04:34 | gibi | mriedem: you are right CoreFilter is not enabled by default in Ocata. I will quickly backport the test locally to Ocata to confirm... | |
| 16:04:36 | jaypipes | dansmith: ^ sorry, I've been in meetings and dealing with russian visa mess all morning and now have a doctor's appt. | |
| 16:04:51 | jaypipes | dansmith: remove the extraneous continue thign | |
| 16:04:55 | dansmith | okay | |
| 16:06:03 | mriedem | dansmith: yeah planned on skipping cells v2 meeting today | |
| 16:06:11 | dansmith | mriedem: okay | |
| 16:06:32 | mriedem | gibi: the fix in ocata would have to be different probably since this code all got refactored in pike | |
| 16:06:42 | mriedem | gibi: but just wanted to make sure it's not an rc1 regression blocker thing | |
| 16:08:18 | dansmith | mriedem: I actually think we probably should enable the cache.. I thought it was already done on the computes, but it's not. The other use on compute is to calculate the rpc pin, which is cached until restart anyway | |
| 16:09:14 | mnaser | danpb cfriesen -- i'm now getting "Not enough available CPUs to schedule instance. Oversubscription is not possible with pinned instances. Required: 1, actual: 0" | |
| 16:09:28 | mnaser | HOST_CELL: NUMACell(cpu_usage=14,cpuset=set([2,4,6,8,10,12,14,18,20,22,24,26,28,30]),id=0,memory=196562,memory_usage=57344,mempages=[NUMAPagesTopology,NUMAPagesTopology],pinned_cpus=set([2,4,6,8,10,12,14,18,20,22,24,26,28,30]),siblings=[set([8,24]),set([2,18]),set([10,26]),set([12,28]),set([6,22]),set([14,30]),set([4,20])]) | |
| 16:11:34 | danpb | where's your other numa cell | |
| 16:11:44 | danpb | that just shows the first cell | |
| 16:12:16 | mnaser | i dont know why its not listed. i added LOG.debug("HOST_CELL: %s" % host_cell) to _numa_fit_instance_cell_with_pinning | |
| 16:12:40 | cfriesen | mnaser: I think I know what's going on...you now have fewer pCPUs in one node than the other, and you're asking for 2-node guests | |
| 16:12:42 | mnaser | let me get you | |
| 16:12:43 | mnaser | both host cell output | |
| 16:12:57 | mnaser | oh | |
| 16:13:01 | mnaser | i think you're right | |
| 16:13:11 | danpb | cfriesen: yep makes sense | |
| 16:13:47 | mnaser | let me switch things back to how they were and get the output of both host cells | |
| 16:14:02 | danpb | mnaser: why do you want the guests to have multiple virtual numa cells ? | |
| 16:14:46 | cfriesen | mriedem: for a reno for https://review.openstack.org/#/c/491854/ would we want to describe the removal of the two default filters in "features", "upgrade", or "other"? | |
| 16:14:52 | danpb | its generally not something you'd do unless guest memory exceeds the amount available in a single host, or need to consume say, PCI devices from separate nods at the same time | |
| 16:15:08 | mnaser | danpb: still experimenting but the idea was more efficent use of hardware. i have 120x 1gb large pages which means that i'll end up with 60gb on each numanode, if i put 8gb sized instances only, i'll end up with 4096mb in each numa node that's unused | |
| 16:15:13 | cfriesen | danpb: or you want increased memory bandwidth | |
| 16:15:54 | danpb | cfriesen: that's only increased if your guest workload avoids cross-node traffic | |
| 16:16:01 | cfriesen | danpb: agreed | |
| 16:16:01 | mnaser | by doing this, it splits 4096 into each numanode and i can fill the freepages .. but it's not something that's set in stone fully | |
| 16:16:11 | danpb | cfriesen: otherwise you'd actually decrease throughput | |
| 16:16:22 | cfriesen | mnaser: the downside is that your guests need to be able to avoid cross-numa traffic, which makes guest coding trickier | |
| 16:16:24 | danpb | by having the guest all contend on the cross-node memory bus | |
| 16:16:48 | cfriesen | danpb: yes, it'd be a special-case | |
| 16:17:00 | mnaser | yeah i was thinking the linux kernel already does an ok job handling multiple numa nodes | |
| 16:17:01 | danpb | IOW unless your guest app is intelligent you want to avoid multiple numa nods | |
| 16:17:10 | mnaser | i see | |
| 16:17:28 | cfriesen | mnaser: the kernel does, but not all userspace code does. | |
| 16:17:38 | mriedem | dansmith: ack - gonna be afk for about an hour | |
| 16:17:40 | mnaser | also the only other tradeoff that comes with this is that the threads are not shared which is not ideal | |
| 16:17:46 | mnaser | so that was something i didnt like about doing this | |
| 16:17:49 | mriedem | can't be worst dad of the summer 2 days in a row | |
| 16:17:57 | cfriesen | mnaser: so if you've got a single large app that wants most of that 8GB... | |