| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-09 | |||
| 16:04:34 | gibi | mriedem: you are right CoreFilter is not enabled by default in Ocata. I will quickly backport the test locally to Ocata to confirm... | |
| 16:04:36 | jaypipes | dansmith: ^ sorry, I've been in meetings and dealing with russian visa mess all morning and now have a doctor's appt. | |
| 16:04:51 | jaypipes | dansmith: remove the extraneous continue thign | |
| 16:04:55 | dansmith | okay | |
| 16:06:03 | mriedem | dansmith: yeah planned on skipping cells v2 meeting today | |
| 16:06:11 | dansmith | mriedem: okay | |
| 16:06:32 | mriedem | gibi: the fix in ocata would have to be different probably since this code all got refactored in pike | |
| 16:06:42 | mriedem | gibi: but just wanted to make sure it's not an rc1 regression blocker thing | |
| 16:08:18 | dansmith | mriedem: I actually think we probably should enable the cache.. I thought it was already done on the computes, but it's not. The other use on compute is to calculate the rpc pin, which is cached until restart anyway | |
| 16:09:14 | mnaser | danpb cfriesen -- i'm now getting "Not enough available CPUs to schedule instance. Oversubscription is not possible with pinned instances. Required: 1, actual: 0" | |
| 16:09:28 | mnaser | HOST_CELL: NUMACell(cpu_usage=14,cpuset=set([2,4,6,8,10,12,14,18,20,22,24,26,28,30]),id=0,memory=196562,memory_usage=57344,mempages=[NUMAPagesTopology,NUMAPagesTopology],pinned_cpus=set([2,4,6,8,10,12,14,18,20,22,24,26,28,30]),siblings=[set([8,24]),set([2,18]),set([10,26]),set([12,28]),set([6,22]),set([14,30]),set([4,20])]) | |
| 16:11:34 | danpb | where's your other numa cell | |
| 16:11:44 | danpb | that just shows the first cell | |
| 16:12:16 | mnaser | i dont know why its not listed. i added LOG.debug("HOST_CELL: %s" % host_cell) to _numa_fit_instance_cell_with_pinning | |
| 16:12:40 | cfriesen | mnaser: I think I know what's going on...you now have fewer pCPUs in one node than the other, and you're asking for 2-node guests | |
| 16:12:42 | mnaser | let me get you | |
| 16:12:43 | mnaser | both host cell output | |
| 16:12:57 | mnaser | oh | |
| 16:13:01 | mnaser | i think you're right | |
| 16:13:11 | danpb | cfriesen: yep makes sense | |
| 16:13:47 | mnaser | let me switch things back to how they were and get the output of both host cells | |
| 16:14:02 | danpb | mnaser: why do you want the guests to have multiple virtual numa cells ? | |
| 16:14:46 | cfriesen | mriedem: for a reno for https://review.openstack.org/#/c/491854/ would we want to describe the removal of the two default filters in "features", "upgrade", or "other"? | |
| 16:14:52 | danpb | its generally not something you'd do unless guest memory exceeds the amount available in a single host, or need to consume say, PCI devices from separate nods at the same time | |
| 16:15:08 | mnaser | danpb: still experimenting but the idea was more efficent use of hardware. i have 120x 1gb large pages which means that i'll end up with 60gb on each numanode, if i put 8gb sized instances only, i'll end up with 4096mb in each numa node that's unused | |
| 16:15:13 | cfriesen | danpb: or you want increased memory bandwidth | |
| 16:15:54 | danpb | cfriesen: that's only increased if your guest workload avoids cross-node traffic | |
| 16:16:01 | mnaser | by doing this, it splits 4096 into each numanode and i can fill the freepages .. but it's not something that's set in stone fully | |
| 16:16:01 | cfriesen | danpb: agreed | |
| 16:16:11 | danpb | cfriesen: otherwise you'd actually decrease throughput | |
| 16:16:22 | cfriesen | mnaser: the downside is that your guests need to be able to avoid cross-numa traffic, which makes guest coding trickier | |
| 16:16:24 | danpb | by having the guest all contend on the cross-node memory bus | |
| 16:16:48 | cfriesen | danpb: yes, it'd be a special-case | |
| 16:17:00 | mnaser | yeah i was thinking the linux kernel already does an ok job handling multiple numa nodes | |
| 16:17:01 | danpb | IOW unless your guest app is intelligent you want to avoid multiple numa nods | |
| 16:17:10 | mnaser | i see | |
| 16:17:28 | cfriesen | mnaser: the kernel does, but not all userspace code does. | |
| 16:17:38 | mriedem | dansmith: ack - gonna be afk for about an hour | |
| 16:17:40 | mnaser | also the only other tradeoff that comes with this is that the threads are not shared which is not ideal | |
| 16:17:46 | mnaser | so that was something i didnt like about doing this | |
| 16:17:49 | mriedem | can't be worst dad of the summer 2 days in a row | |
| 16:17:57 | cfriesen | mnaser: so if you've got a single large app that wants most of that 8GB... | |
| 16:18:16 | dansmith | mriedem: okay I'll push this up with those changes and I'll be gone before jenkins gets to it | |
| 16:18:23 | mnaser | for completition sake btw, this is both host cells | |
| 16:18:24 | dansmith | mriedem: jaypipes so baton back to you at that point | |
| 16:18:45 | mnaser | http://paste.openstack.org/show/617957/ | |
| 16:19:26 | mnaser | there are 15 sets there | |
| 16:19:37 | mnaser | let me see what it was when we had 0-1 only | |
| 16:20:17 | cdent | dansmith: couple questions about expected discover_hosts behavior: is it supposed to be idempotent (run it again and again, it’s okay)? It is supposed to cope if two different processes run it at the same time? | |
| 16:20:24 | cfriesen | mnaser: but only 14 pairs of pCPUs from different numa nodes | |
| 16:20:43 | mnaser | yeah, thats why it failed now, but in the original case | |
| 16:20:44 | mnaser | strangely enough | |
| 16:20:51 | dansmith | cdent: yeah, it's expected to just run it over and over again, from cron even | |
| 16:20:51 | mnaser | i only see one output of numacell, not two | |
| 16:21:15 | dansmith | cdent: you might get a failure if you run multiples and one loses the race to insert the record, but other than that it should be fine to run multiple threads of it | |
| 16:21:41 | cdent | dansmith: thanks that’s what I was hoping/expecting but wanted to confirm | |
| 16:23:51 | mnaser | im going to try setting hw:numa_nodes=1 and packing the server and seeing what happens (i should be able to get 14 at least) | |
| 16:25:47 | gibi | mriedem: ocata is not affected by the bug: https://github.com/openstack/nova/blob/stable/ocata/nova/scheduler/filter_scheduler.py#L188 | |
| 16:26:19 | gibi | mriedem: I mean bug 1708637 | |
| 16:26:21 | openstack | bug 1708637 in OpenStack Compute (nova) "nova does not properly claim resources when server resized to a too big flavor" [High,In progress] https://launchpad.net/bugs/1708637 - Assigned to Balazs Gibizer (balazs-gibizer) | |
| 16:26:33 | cfriesen | mnaser: do you actually need 8GB? if you don't actually need it all, you could drop to 2MB hugepages and divide up the memory evenly with less waste. For most things 1GB pages don't give that big of a boost. | |
| 16:27:35 | mnaser | cfriesen have you had experience with it? i just figured that if i can have 1gb pages, it would be better than 2mb pages but there isn't much substance to it other than 'it seems right' | |
| 16:27:42 | mnaser | dropping to 2mb would obviously make life much easier | |
| 16:27:54 | cdent | mriedem: yeah, we’ve been pretty inconsistent about which 400s are document in the placement-api-ref. anything common like “yo, not here” and “hey, you violated schema” has frequently been dropped. I’ve not been too strict on my reviews of that stuff except where a 4xx has some particular weird sense | |
| 16:28:39 | cdent | the current tooling doesn’t really have as much support for handling error responses as we might want, that’s probably something we could and should address later. I think once the whole thing is documented we’ll be able to tune it as a whole, better | |
| 16:29:14 | mnaser | i just tried to create 14x 2vcpu/8gb memory and that went ok, then 1x 1vcpu/4gb memory and that was okay, so now i have 4gb free large pages.. trying to create 1vcpu/4gb - Host does not support requested memory pagesize. Requested: 1048576 kB | |
| 16:30:42 | mnaser | looks like it was assigned core 21, which is on node1 .. but my free 4096 is on node0 | |
| 16:30:53 | mnaser | i guess this happened because im still using the same thread pair, probably wouldnt happen again if i reserve 0-1 for the os | |
| 16:31:54 | cfriesen | mnaser: we've done some testing. there are some cases where it makes a noticeable difference, but most of the time the difference is not worth the wasted memory. it depends on the guest | |
| 16:33:02 | mnaser | cfriesen all the guests are going to be multiples of 1gb in memory so 120gb in large pages in 1gb or in 2mb .. wouldn't make much difference, no? it looks like the # of freepages is split across both numa nodes | |
| 16:34:59 | cfriesen | mnaser: if it must be exact multiples of 1GB, then you may as well use 1GB hugepages. | |
| 16:35:43 | mnaser | cfriesen yeah they're all multiples of 1gb.. anyways ill try going back to 0-1 and seeing if i can fully populate the server | |
| 16:35:55 | mnaser | and then after that ill file a bug regarding that sibling_set issue | |
| 16:44:33 | cdent | mriedem: is this still relevant or has other stuff killed it: https://review.openstack.org/#/c/488187/ | |
| 17:13:05 | mnaser | reserving 0-1 allows me to start 14 2vcpu/8gb, but refuses to let me start anymore with that same sibling_sets bug | |
| 17:14:19 | openstackgerrit | Chris Friesen proposed openstack/nova master: Remove ram/disk sched filters from default list https://review.openstack.org/491854 | |
| 17:17:16 | openstackgerrit | Chris Dent proposed openstack/nova master: replace chance with filter scheduler in func tests https://review.openstack.org/491529 | |
| 17:18:21 | mnaser | cfriesen: doing more debugging, i believe i'm onto something -- HOST_TOPOLOGY: NUMATopology(cells=[NUMACell(UNKNOWN),NUMACell(1)]) | |
| 17:18:42 | mnaser | for some reason the first numacell is unknown..? i'll have to check that | |
| 17:23:23 | melwitt | dansmith: +1 to skipping cells meeting | |
| 17:29:17 | mriedem | cdent: not sure, but not something we need to care about for rc1 | |
| 17:29:52 | cdent | mriedem: ‘k, will circle back round to that after we get the main things settled | |
| 17:31:20 | openstackgerrit | Dan Smith proposed openstack/nova master: Resource tracker compatibility with Ocata and Pike https://review.openstack.org/491012 | |
| 17:31:30 | dansmith | mriedem: jaypipes ^ | |
| 17:33:00 | mriedem | cfriesen: commented in your change | |
| 17:33:16 | mriedem | cfriesen: if you're using the caching scheduler, you still rely on those filters since the caching scheduler doesn't use placement | |
| 17:33:37 | dansmith | mriedem: did we ever push up that deprecation for those schedulers? | |
| 17:33:38 | mriedem | but the caching scheduler isn't the default filter driver, the filter_scheduler is, so i think it's fine for the default enabled filters to match the default scheduler driver | |
| 17:33:43 | mriedem | dansmith: i didn't see one | |
| 17:33:45 | dansmith | sounded like sdague was going to but I don't know that he did | |
| 17:34:04 | dansmith | mriedem: is it enough to put something in the config help text or do you want to log something? | |
| 17:34:48 | mriedem | i think we'd need both | |
| 17:35:09 | mnaser | cfriesen i'm onto something, with a totally empty hypervisor -- i see this -- AVAILABLE_SIBLINGS: [CoercedSet([1, 17]), CoercedSet([31, 15]), CoercedSet([23, 7]), CoercedSet([13, 29]), CoercedSet([27, 11]), CoercedSet([19, 3]), CoercedSet([9, 25]), CoercedSet([5, 21])] | |
| 17:35:26 | dansmith | mriedem: okay | |
| 17:35:27 | cdent | dansmith: was the resolution on the question of what to do when we get to the end of that set of conditionals where jay had “literally no idea what to do” to do nothing? | |
| 17:35:32 | mnaser | [1, 17] shouldn't be there, it should just be 17, because vcpu_pin_set is 2-31 | |
| 17:35:32 | mriedem | dansmith: plus reno of course | |
| 17:35:43 | mnaser | i dont think that codebase takes vcpu_pin_set into consideration | |
| 17:35:49 | dansmith | cdent: yes | |
| 17:35:55 | cdent | ✔ | |