| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-09 | |||
| 15:52:57 | dansmith | I really don't know anything about this stuff | |
| 15:54:53 | jaypipes | mriedem, dansmith: I will be afk for 2 hours, just FYI. | |
| 15:54:55 | cfriesen | mnaser: so are 16/17 the siblings of 0/1 which are not included in vcpu_pin_set? | |
| 15:55:08 | openstackgerrit | Matt Riedemann proposed openstack/nova master: [placement] Add api-ref for usages https://review.openstack.org/480563 | |
| 15:55:11 | artom | About the fix, at any rate. | |
| 15:55:13 | mnaser | cfriesen nope, 0/1 are the ones which are not included | |
| 15:55:26 | mnaser | so our vcpu_pin_set is 2-31 | |
| 15:55:30 | mnaser | (32 core machine) | |
| 15:55:33 | cfriesen | mnaser: yes, but which are the HT siblings of 0/1? | |
| 15:56:15 | danpb | hmm, yes, it seems numa node 0 has even numbered cpus, and node 1 has odd numbered cpus | |
| 15:56:19 | mnaser | cfriesen sorry, not sure i follow, but here's lscpu output if that helps? http://paste.openstack.org/show/617955/ | |
| 15:56:34 | danpb | so excluding 0-1, kills 1 cpu each host numa node, instead of 2 cpus | |
| 15:56:35 | cfriesen | mnaser: based on the pattern above, the HT siblings of 0/1 would be 16/17....I'm wondering if that's confusing the logic in hardware.py | |
| 15:56:42 | mnaser | but i think that they are | |
| 15:57:19 | cfriesen | mnaswer: running "virsh capabilities" will show sibling info | |
| 15:57:26 | danpb | so that in turn kills a entire sibling pair from each numa node | |
| 15:57:37 | danpb | effectively making 4 cpus unavailable | |
| 15:57:53 | cfriesen | danpb: theoretically it shouldn't kill the sibling pair though...we should have two pCPUs with no siblings | |
| 15:58:03 | mnaser | here is output of virsh caps. http://paste.openstack.org/show/617956/ | |
| 15:58:06 | cfriesen | but maybe the logic can't handle a mix of sibs and no sibs | |
| 15:58:15 | mriedem | someone want to approve this? https://review.openstack.org/#/c/491855/ | |
| 15:58:18 | danpb | yeah that wouldn't be surprising | |
| 15:58:46 | mnaser | so am i better off reserving 2 siblings and maybe that might get rid of the 'confusion' ? | |
| 15:58:53 | cfriesen | mnaswer: seems likely | |
| 15:58:55 | mnaser | so instead of 0,1 .. do 0,16 instead? | |
| 15:59:02 | danpb | yeah | |
| 15:59:53 | mnaser | let me test that theory out | |
| 16:00:03 | openstackgerrit | Ed Leafe proposed openstack/nova master: Handle addition of new nodes/instances in ironic flavor migration https://review.openstack.org/487954 | |
| 16:00:12 | edleafe | dansmith: dtantsur: ^^ updated to remove the node RC change handling | |
| 16:00:37 | dtantsur | awesome! I'm just out of a meeting, will do the ironic part now | |
| 16:00:50 | mnaser | so my vcpu_pin_set should technically be 1-15,17-31 in that case, correct? | |
| 16:00:53 | mnaser | leaving 0,16 for the OS | |
| 16:00:53 | mriedem | gibi: wasn't that also a problem in ocata then? CoreFilter isn't enabled by default | |
| 16:00:59 | cfriesen | mnaswer: either way, please raise a bug for this issue, it'd be nice to handle it properly | |
| 16:01:02 | danpb | mnaser: yeah | |
| 16:01:11 | mnaser | cfriesen agreed | |
| 16:01:15 | mnaser | ok, lets try this out | |
| 16:01:35 | mnaser | ill reboot with the new updates isolcpus= values so itll take me a few minutes and report | |
| 16:02:31 | dansmith | mriedem: melwitt: I'm going to disappear suddenly somewhere in the hour of the cells meeting.. assume we were on track to punt anyway, but.. that okay? | |
| 16:03:57 | dansmith | jaypipes: so what's the deal with that last patch? do you want me to update it and remove the extraneous continue? or log something there or what? | |
| 16:04:15 | jaypipes | please, yes, go for it. | |
| 16:04:32 | dansmith | ...which? | |
| 16:04:34 | gibi | mriedem: you are right CoreFilter is not enabled by default in Ocata. I will quickly backport the test locally to Ocata to confirm... | |
| 16:04:36 | jaypipes | dansmith: ^ sorry, I've been in meetings and dealing with russian visa mess all morning and now have a doctor's appt. | |
| 16:04:51 | jaypipes | dansmith: remove the extraneous continue thign | |
| 16:04:55 | dansmith | okay | |
| 16:06:03 | mriedem | dansmith: yeah planned on skipping cells v2 meeting today | |
| 16:06:11 | dansmith | mriedem: okay | |
| 16:06:32 | mriedem | gibi: the fix in ocata would have to be different probably since this code all got refactored in pike | |
| 16:06:42 | mriedem | gibi: but just wanted to make sure it's not an rc1 regression blocker thing | |
| 16:08:18 | dansmith | mriedem: I actually think we probably should enable the cache.. I thought it was already done on the computes, but it's not. The other use on compute is to calculate the rpc pin, which is cached until restart anyway | |
| 16:09:14 | mnaser | danpb cfriesen -- i'm now getting "Not enough available CPUs to schedule instance. Oversubscription is not possible with pinned instances. Required: 1, actual: 0" | |
| 16:09:28 | mnaser | HOST_CELL: NUMACell(cpu_usage=14,cpuset=set([2,4,6,8,10,12,14,18,20,22,24,26,28,30]),id=0,memory=196562,memory_usage=57344,mempages=[NUMAPagesTopology,NUMAPagesTopology],pinned_cpus=set([2,4,6,8,10,12,14,18,20,22,24,26,28,30]),siblings=[set([8,24]),set([2,18]),set([10,26]),set([12,28]),set([6,22]),set([14,30]),set([4,20])]) | |
| 16:11:34 | danpb | where's your other numa cell | |
| 16:11:44 | danpb | that just shows the first cell | |
| 16:12:16 | mnaser | i dont know why its not listed. i added LOG.debug("HOST_CELL: %s" % host_cell) to _numa_fit_instance_cell_with_pinning | |
| 16:12:40 | cfriesen | mnaser: I think I know what's going on...you now have fewer pCPUs in one node than the other, and you're asking for 2-node guests | |
| 16:12:42 | mnaser | let me get you | |
| 16:12:43 | mnaser | both host cell output | |
| 16:12:57 | mnaser | oh | |
| 16:13:01 | mnaser | i think you're right | |
| 16:13:11 | danpb | cfriesen: yep makes sense | |
| 16:13:47 | mnaser | let me switch things back to how they were and get the output of both host cells | |
| 16:14:02 | danpb | mnaser: why do you want the guests to have multiple virtual numa cells ? | |
| 16:14:46 | cfriesen | mriedem: for a reno for https://review.openstack.org/#/c/491854/ would we want to describe the removal of the two default filters in "features", "upgrade", or "other"? | |
| 16:14:52 | danpb | its generally not something you'd do unless guest memory exceeds the amount available in a single host, or need to consume say, PCI devices from separate nods at the same time | |
| 16:15:08 | mnaser | danpb: still experimenting but the idea was more efficent use of hardware. i have 120x 1gb large pages which means that i'll end up with 60gb on each numanode, if i put 8gb sized instances only, i'll end up with 4096mb in each numa node that's unused | |
| 16:15:13 | cfriesen | danpb: or you want increased memory bandwidth | |
| 16:15:54 | danpb | cfriesen: that's only increased if your guest workload avoids cross-node traffic | |
| 16:16:01 | mnaser | by doing this, it splits 4096 into each numanode and i can fill the freepages .. but it's not something that's set in stone fully | |
| 16:16:01 | cfriesen | danpb: agreed | |
| 16:16:11 | danpb | cfriesen: otherwise you'd actually decrease throughput | |
| 16:16:22 | cfriesen | mnaser: the downside is that your guests need to be able to avoid cross-numa traffic, which makes guest coding trickier | |
| 16:16:24 | danpb | by having the guest all contend on the cross-node memory bus | |
| 16:16:48 | cfriesen | danpb: yes, it'd be a special-case | |
| 16:17:00 | mnaser | yeah i was thinking the linux kernel already does an ok job handling multiple numa nodes | |
| 16:17:01 | danpb | IOW unless your guest app is intelligent you want to avoid multiple numa nods | |
| 16:17:10 | mnaser | i see | |
| 16:17:28 | cfriesen | mnaser: the kernel does, but not all userspace code does. | |
| 16:17:38 | mriedem | dansmith: ack - gonna be afk for about an hour | |
| 16:17:40 | mnaser | also the only other tradeoff that comes with this is that the threads are not shared which is not ideal | |
| 16:17:46 | mnaser | so that was something i didnt like about doing this | |
| 16:17:49 | mriedem | can't be worst dad of the summer 2 days in a row | |
| 16:17:57 | cfriesen | mnaser: so if you've got a single large app that wants most of that 8GB... | |
| 16:18:16 | dansmith | mriedem: okay I'll push this up with those changes and I'll be gone before jenkins gets to it | |
| 16:18:23 | mnaser | for completition sake btw, this is both host cells | |
| 16:18:24 | dansmith | mriedem: jaypipes so baton back to you at that point | |
| 16:18:45 | mnaser | http://paste.openstack.org/show/617957/ | |
| 16:19:26 | mnaser | there are 15 sets there | |
| 16:19:37 | mnaser | let me see what it was when we had 0-1 only | |
| 16:20:17 | cdent | dansmith: couple questions about expected discover_hosts behavior: is it supposed to be idempotent (run it again and again, it’s okay)? It is supposed to cope if two different processes run it at the same time? | |
| 16:20:24 | cfriesen | mnaser: but only 14 pairs of pCPUs from different numa nodes | |
| 16:20:43 | mnaser | yeah, thats why it failed now, but in the original case | |
| 16:20:44 | mnaser | strangely enough | |
| 16:20:51 | dansmith | cdent: yeah, it's expected to just run it over and over again, from cron even | |
| 16:20:51 | mnaser | i only see one output of numacell, not two | |
| 16:21:15 | dansmith | cdent: you might get a failure if you run multiples and one loses the race to insert the record, but other than that it should be fine to run multiple threads of it | |
| 16:21:41 | cdent | dansmith: thanks that’s what I was hoping/expecting but wanted to confirm | |
| 16:23:51 | mnaser | im going to try setting hw:numa_nodes=1 and packing the server and seeing what happens (i should be able to get 14 at least) | |
| 16:25:47 | gibi | mriedem: ocata is not affected by the bug: https://github.com/openstack/nova/blob/stable/ocata/nova/scheduler/filter_scheduler.py#L188 | |