| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-09 | |||
| 15:55:13 | mnaser | cfriesen nope, 0/1 are the ones which are not included | |
| 15:55:26 | mnaser | so our vcpu_pin_set is 2-31 | |
| 15:55:30 | mnaser | (32 core machine) | |
| 15:55:33 | cfriesen | mnaser: yes, but which are the HT siblings of 0/1? | |
| 15:56:15 | danpb | hmm, yes, it seems numa node 0 has even numbered cpus, and node 1 has odd numbered cpus | |
| 15:56:19 | mnaser | cfriesen sorry, not sure i follow, but here's lscpu output if that helps? http://paste.openstack.org/show/617955/ | |
| 15:56:34 | danpb | so excluding 0-1, kills 1 cpu each host numa node, instead of 2 cpus | |
| 15:56:35 | cfriesen | mnaser: based on the pattern above, the HT siblings of 0/1 would be 16/17....I'm wondering if that's confusing the logic in hardware.py | |
| 15:56:42 | mnaser | but i think that they are | |
| 15:57:19 | cfriesen | mnaswer: running "virsh capabilities" will show sibling info | |
| 15:57:26 | danpb | so that in turn kills a entire sibling pair from each numa node | |
| 15:57:37 | danpb | effectively making 4 cpus unavailable | |
| 15:57:53 | cfriesen | danpb: theoretically it shouldn't kill the sibling pair though...we should have two pCPUs with no siblings | |
| 15:58:03 | mnaser | here is output of virsh caps. http://paste.openstack.org/show/617956/ | |
| 15:58:06 | cfriesen | but maybe the logic can't handle a mix of sibs and no sibs | |
| 15:58:15 | mriedem | someone want to approve this? https://review.openstack.org/#/c/491855/ | |
| 15:58:18 | danpb | yeah that wouldn't be surprising | |
| 15:58:46 | mnaser | so am i better off reserving 2 siblings and maybe that might get rid of the 'confusion' ? | |
| 15:58:53 | cfriesen | mnaswer: seems likely | |
| 15:58:55 | mnaser | so instead of 0,1 .. do 0,16 instead? | |
| 15:59:02 | danpb | yeah | |
| 15:59:53 | mnaser | let me test that theory out | |
| 16:00:03 | openstackgerrit | Ed Leafe proposed openstack/nova master: Handle addition of new nodes/instances in ironic flavor migration https://review.openstack.org/487954 | |
| 16:00:12 | edleafe | dansmith: dtantsur: ^^ updated to remove the node RC change handling | |
| 16:00:37 | dtantsur | awesome! I'm just out of a meeting, will do the ironic part now | |
| 16:00:50 | mnaser | so my vcpu_pin_set should technically be 1-15,17-31 in that case, correct? | |
| 16:00:53 | mnaser | leaving 0,16 for the OS | |
| 16:00:53 | mriedem | gibi: wasn't that also a problem in ocata then? CoreFilter isn't enabled by default | |
| 16:00:59 | cfriesen | mnaswer: either way, please raise a bug for this issue, it'd be nice to handle it properly | |
| 16:01:02 | danpb | mnaser: yeah | |
| 16:01:11 | mnaser | cfriesen agreed | |
| 16:01:15 | mnaser | ok, lets try this out | |
| 16:01:35 | mnaser | ill reboot with the new updates isolcpus= values so itll take me a few minutes and report | |
| 16:02:31 | dansmith | mriedem: melwitt: I'm going to disappear suddenly somewhere in the hour of the cells meeting.. assume we were on track to punt anyway, but.. that okay? | |
| 16:03:57 | dansmith | jaypipes: so what's the deal with that last patch? do you want me to update it and remove the extraneous continue? or log something there or what? | |
| 16:04:15 | jaypipes | please, yes, go for it. | |
| 16:04:32 | dansmith | ...which? | |
| 16:04:34 | gibi | mriedem: you are right CoreFilter is not enabled by default in Ocata. I will quickly backport the test locally to Ocata to confirm... | |
| 16:04:36 | jaypipes | dansmith: ^ sorry, I've been in meetings and dealing with russian visa mess all morning and now have a doctor's appt. | |
| 16:04:51 | jaypipes | dansmith: remove the extraneous continue thign | |
| 16:04:55 | dansmith | okay | |
| 16:06:03 | mriedem | dansmith: yeah planned on skipping cells v2 meeting today | |
| 16:06:11 | dansmith | mriedem: okay | |
| 16:06:32 | mriedem | gibi: the fix in ocata would have to be different probably since this code all got refactored in pike | |
| 16:06:42 | mriedem | gibi: but just wanted to make sure it's not an rc1 regression blocker thing | |
| 16:08:18 | dansmith | mriedem: I actually think we probably should enable the cache.. I thought it was already done on the computes, but it's not. The other use on compute is to calculate the rpc pin, which is cached until restart anyway | |
| 16:09:14 | mnaser | danpb cfriesen -- i'm now getting "Not enough available CPUs to schedule instance. Oversubscription is not possible with pinned instances. Required: 1, actual: 0" | |
| 16:09:28 | mnaser | HOST_CELL: NUMACell(cpu_usage=14,cpuset=set([2,4,6,8,10,12,14,18,20,22,24,26,28,30]),id=0,memory=196562,memory_usage=57344,mempages=[NUMAPagesTopology,NUMAPagesTopology],pinned_cpus=set([2,4,6,8,10,12,14,18,20,22,24,26,28,30]),siblings=[set([8,24]),set([2,18]),set([10,26]),set([12,28]),set([6,22]),set([14,30]),set([4,20])]) | |
| 16:11:34 | danpb | where's your other numa cell | |
| 16:11:44 | danpb | that just shows the first cell | |
| 16:12:16 | mnaser | i dont know why its not listed. i added LOG.debug("HOST_CELL: %s" % host_cell) to _numa_fit_instance_cell_with_pinning | |
| 16:12:40 | cfriesen | mnaser: I think I know what's going on...you now have fewer pCPUs in one node than the other, and you're asking for 2-node guests | |
| 16:12:42 | mnaser | let me get you | |
| 16:12:43 | mnaser | both host cell output | |
| 16:12:57 | mnaser | oh | |
| 16:13:01 | mnaser | i think you're right | |
| 16:13:11 | danpb | cfriesen: yep makes sense | |
| 16:13:47 | mnaser | let me switch things back to how they were and get the output of both host cells | |
| 16:14:02 | danpb | mnaser: why do you want the guests to have multiple virtual numa cells ? | |
| 16:14:46 | cfriesen | mriedem: for a reno for https://review.openstack.org/#/c/491854/ would we want to describe the removal of the two default filters in "features", "upgrade", or "other"? | |
| 16:14:52 | danpb | its generally not something you'd do unless guest memory exceeds the amount available in a single host, or need to consume say, PCI devices from separate nods at the same time | |
| 16:15:08 | mnaser | danpb: still experimenting but the idea was more efficent use of hardware. i have 120x 1gb large pages which means that i'll end up with 60gb on each numanode, if i put 8gb sized instances only, i'll end up with 4096mb in each numa node that's unused | |
| 16:15:13 | cfriesen | danpb: or you want increased memory bandwidth | |
| 16:15:54 | danpb | cfriesen: that's only increased if your guest workload avoids cross-node traffic | |
| 16:16:01 | mnaser | by doing this, it splits 4096 into each numanode and i can fill the freepages .. but it's not something that's set in stone fully | |
| 16:16:01 | cfriesen | danpb: agreed | |
| 16:16:11 | danpb | cfriesen: otherwise you'd actually decrease throughput | |
| 16:16:22 | cfriesen | mnaser: the downside is that your guests need to be able to avoid cross-numa traffic, which makes guest coding trickier | |
| 16:16:24 | danpb | by having the guest all contend on the cross-node memory bus | |
| 16:16:48 | cfriesen | danpb: yes, it'd be a special-case | |
| 16:17:00 | mnaser | yeah i was thinking the linux kernel already does an ok job handling multiple numa nodes | |
| 16:17:01 | danpb | IOW unless your guest app is intelligent you want to avoid multiple numa nods | |
| 16:17:10 | mnaser | i see | |
| 16:17:28 | cfriesen | mnaser: the kernel does, but not all userspace code does. | |
| 16:17:38 | mriedem | dansmith: ack - gonna be afk for about an hour | |
| 16:17:40 | mnaser | also the only other tradeoff that comes with this is that the threads are not shared which is not ideal | |
| 16:17:46 | mnaser | so that was something i didnt like about doing this | |
| 16:17:49 | mriedem | can't be worst dad of the summer 2 days in a row | |
| 16:17:57 | cfriesen | mnaser: so if you've got a single large app that wants most of that 8GB... | |
| 16:18:16 | dansmith | mriedem: okay I'll push this up with those changes and I'll be gone before jenkins gets to it | |
| 16:18:23 | mnaser | for completition sake btw, this is both host cells | |
| 16:18:24 | dansmith | mriedem: jaypipes so baton back to you at that point | |
| 16:18:45 | mnaser | http://paste.openstack.org/show/617957/ | |
| 16:19:26 | mnaser | there are 15 sets there | |
| 16:19:37 | mnaser | let me see what it was when we had 0-1 only | |
| 16:20:17 | cdent | dansmith: couple questions about expected discover_hosts behavior: is it supposed to be idempotent (run it again and again, it’s okay)? It is supposed to cope if two different processes run it at the same time? | |
| 16:20:24 | cfriesen | mnaser: but only 14 pairs of pCPUs from different numa nodes | |
| 16:20:43 | mnaser | yeah, thats why it failed now, but in the original case | |
| 16:20:44 | mnaser | strangely enough | |
| 16:20:51 | dansmith | cdent: yeah, it's expected to just run it over and over again, from cron even | |
| 16:20:51 | mnaser | i only see one output of numacell, not two | |
| 16:21:15 | dansmith | cdent: you might get a failure if you run multiples and one loses the race to insert the record, but other than that it should be fine to run multiple threads of it | |
| 16:21:41 | cdent | dansmith: thanks that’s what I was hoping/expecting but wanted to confirm | |
| 16:23:51 | mnaser | im going to try setting hw:numa_nodes=1 and packing the server and seeing what happens (i should be able to get 14 at least) | |
| 16:25:47 | gibi | mriedem: ocata is not affected by the bug: https://github.com/openstack/nova/blob/stable/ocata/nova/scheduler/filter_scheduler.py#L188 | |
| 16:26:19 | gibi | mriedem: I mean bug 1708637 | |
| 16:26:21 | openstack | bug 1708637 in OpenStack Compute (nova) "nova does not properly claim resources when server resized to a too big flavor" [High,In progress] https://launchpad.net/bugs/1708637 - Assigned to Balazs Gibizer (balazs-gibizer) | |
| 16:26:33 | cfriesen | mnaser: do you actually need 8GB? if you don't actually need it all, you could drop to 2MB hugepages and divide up the memory evenly with less waste. For most things 1GB pages don't give that big of a boost. | |
| 16:27:35 | mnaser | cfriesen have you had experience with it? i just figured that if i can have 1gb pages, it would be better than 2mb pages but there isn't much substance to it other than 'it seems right' | |
| 16:27:42 | mnaser | dropping to 2mb would obviously make life much easier | |