| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-09 | |||
| 16:13:47 | mnaser | let me switch things back to how they were and get the output of both host cells | |
| 16:14:02 | danpb | mnaser: why do you want the guests to have multiple virtual numa cells ? | |
| 16:14:46 | cfriesen | mriedem: for a reno for https://review.openstack.org/#/c/491854/ would we want to describe the removal of the two default filters in "features", "upgrade", or "other"? | |
| 16:14:52 | danpb | its generally not something you'd do unless guest memory exceeds the amount available in a single host, or need to consume say, PCI devices from separate nods at the same time | |
| 16:15:08 | mnaser | danpb: still experimenting but the idea was more efficent use of hardware. i have 120x 1gb large pages which means that i'll end up with 60gb on each numanode, if i put 8gb sized instances only, i'll end up with 4096mb in each numa node that's unused | |
| 16:15:13 | cfriesen | danpb: or you want increased memory bandwidth | |
| 16:15:54 | danpb | cfriesen: that's only increased if your guest workload avoids cross-node traffic | |
| 16:16:01 | cfriesen | danpb: agreed | |
| 16:16:01 | mnaser | by doing this, it splits 4096 into each numanode and i can fill the freepages .. but it's not something that's set in stone fully | |
| 16:16:11 | danpb | cfriesen: otherwise you'd actually decrease throughput | |
| 16:16:22 | cfriesen | mnaser: the downside is that your guests need to be able to avoid cross-numa traffic, which makes guest coding trickier | |
| 16:16:24 | danpb | by having the guest all contend on the cross-node memory bus | |
| 16:16:48 | cfriesen | danpb: yes, it'd be a special-case | |
| 16:17:00 | mnaser | yeah i was thinking the linux kernel already does an ok job handling multiple numa nodes | |
| 16:17:01 | danpb | IOW unless your guest app is intelligent you want to avoid multiple numa nods | |
| 16:17:10 | mnaser | i see | |
| 16:17:28 | cfriesen | mnaser: the kernel does, but not all userspace code does. | |
| 16:17:38 | mriedem | dansmith: ack - gonna be afk for about an hour | |
| 16:17:40 | mnaser | also the only other tradeoff that comes with this is that the threads are not shared which is not ideal | |
| 16:17:46 | mnaser | so that was something i didnt like about doing this | |
| 16:17:49 | mriedem | can't be worst dad of the summer 2 days in a row | |
| 16:17:57 | cfriesen | mnaser: so if you've got a single large app that wants most of that 8GB... | |
| 16:18:16 | dansmith | mriedem: okay I'll push this up with those changes and I'll be gone before jenkins gets to it | |
| 16:18:23 | mnaser | for completition sake btw, this is both host cells | |
| 16:18:24 | dansmith | mriedem: jaypipes so baton back to you at that point | |
| 16:18:45 | mnaser | http://paste.openstack.org/show/617957/ | |
| 16:19:26 | mnaser | there are 15 sets there | |
| 16:19:37 | mnaser | let me see what it was when we had 0-1 only | |
| 16:20:17 | cdent | dansmith: couple questions about expected discover_hosts behavior: is it supposed to be idempotent (run it again and again, it’s okay)? It is supposed to cope if two different processes run it at the same time? | |
| 16:20:24 | cfriesen | mnaser: but only 14 pairs of pCPUs from different numa nodes | |
| 16:20:43 | mnaser | yeah, thats why it failed now, but in the original case | |
| 16:20:44 | mnaser | strangely enough | |
| 16:20:51 | mnaser | i only see one output of numacell, not two | |
| 16:20:51 | dansmith | cdent: yeah, it's expected to just run it over and over again, from cron even | |
| 16:21:15 | dansmith | cdent: you might get a failure if you run multiples and one loses the race to insert the record, but other than that it should be fine to run multiple threads of it | |
| 16:21:41 | cdent | dansmith: thanks that’s what I was hoping/expecting but wanted to confirm | |
| 16:23:51 | mnaser | im going to try setting hw:numa_nodes=1 and packing the server and seeing what happens (i should be able to get 14 at least) | |
| 16:25:47 | gibi | mriedem: ocata is not affected by the bug: https://github.com/openstack/nova/blob/stable/ocata/nova/scheduler/filter_scheduler.py#L188 | |
| 16:26:19 | gibi | mriedem: I mean bug 1708637 | |
| 16:26:21 | openstack | bug 1708637 in OpenStack Compute (nova) "nova does not properly claim resources when server resized to a too big flavor" [High,In progress] https://launchpad.net/bugs/1708637 - Assigned to Balazs Gibizer (balazs-gibizer) | |
| 16:26:33 | cfriesen | mnaser: do you actually need 8GB? if you don't actually need it all, you could drop to 2MB hugepages and divide up the memory evenly with less waste. For most things 1GB pages don't give that big of a boost. | |
| 16:27:35 | mnaser | cfriesen have you had experience with it? i just figured that if i can have 1gb pages, it would be better than 2mb pages but there isn't much substance to it other than 'it seems right' | |
| 16:27:42 | mnaser | dropping to 2mb would obviously make life much easier | |
| 16:27:54 | cdent | mriedem: yeah, we’ve been pretty inconsistent about which 400s are document in the placement-api-ref. anything common like “yo, not here” and “hey, you violated schema” has frequently been dropped. I’ve not been too strict on my reviews of that stuff except where a 4xx has some particular weird sense | |
| 16:28:39 | cdent | the current tooling doesn’t really have as much support for handling error responses as we might want, that’s probably something we could and should address later. I think once the whole thing is documented we’ll be able to tune it as a whole, better | |
| 16:29:14 | mnaser | i just tried to create 14x 2vcpu/8gb memory and that went ok, then 1x 1vcpu/4gb memory and that was okay, so now i have 4gb free large pages.. trying to create 1vcpu/4gb - Host does not support requested memory pagesize. Requested: 1048576 kB | |
| 16:30:42 | mnaser | looks like it was assigned core 21, which is on node1 .. but my free 4096 is on node0 | |
| 16:30:53 | mnaser | i guess this happened because im still using the same thread pair, probably wouldnt happen again if i reserve 0-1 for the os | |
| 16:31:54 | cfriesen | mnaser: we've done some testing. there are some cases where it makes a noticeable difference, but most of the time the difference is not worth the wasted memory. it depends on the guest | |
| 16:33:02 | mnaser | cfriesen all the guests are going to be multiples of 1gb in memory so 120gb in large pages in 1gb or in 2mb .. wouldn't make much difference, no? it looks like the # of freepages is split across both numa nodes | |
| 16:34:59 | cfriesen | mnaser: if it must be exact multiples of 1GB, then you may as well use 1GB hugepages. | |
| 16:35:43 | mnaser | cfriesen yeah they're all multiples of 1gb.. anyways ill try going back to 0-1 and seeing if i can fully populate the server | |
| 16:35:55 | mnaser | and then after that ill file a bug regarding that sibling_set issue | |
| 16:44:33 | cdent | mriedem: is this still relevant or has other stuff killed it: https://review.openstack.org/#/c/488187/ | |
| 17:13:05 | mnaser | reserving 0-1 allows me to start 14 2vcpu/8gb, but refuses to let me start anymore with that same sibling_sets bug | |
| 17:14:19 | openstackgerrit | Chris Friesen proposed openstack/nova master: Remove ram/disk sched filters from default list https://review.openstack.org/491854 | |
| 17:17:16 | openstackgerrit | Chris Dent proposed openstack/nova master: replace chance with filter scheduler in func tests https://review.openstack.org/491529 | |
| 17:18:21 | mnaser | cfriesen: doing more debugging, i believe i'm onto something -- HOST_TOPOLOGY: NUMATopology(cells=[NUMACell(UNKNOWN),NUMACell(1)]) | |
| 17:18:42 | mnaser | for some reason the first numacell is unknown..? i'll have to check that | |
| 17:23:23 | melwitt | dansmith: +1 to skipping cells meeting | |
| 17:29:17 | mriedem | cdent: not sure, but not something we need to care about for rc1 | |
| 17:29:52 | cdent | mriedem: ‘k, will circle back round to that after we get the main things settled | |
| 17:31:20 | openstackgerrit | Dan Smith proposed openstack/nova master: Resource tracker compatibility with Ocata and Pike https://review.openstack.org/491012 | |
| 17:31:30 | dansmith | mriedem: jaypipes ^ | |
| 17:33:00 | mriedem | cfriesen: commented in your change | |
| 17:33:16 | mriedem | cfriesen: if you're using the caching scheduler, you still rely on those filters since the caching scheduler doesn't use placement | |
| 17:33:37 | dansmith | mriedem: did we ever push up that deprecation for those schedulers? | |
| 17:33:38 | mriedem | but the caching scheduler isn't the default filter driver, the filter_scheduler is, so i think it's fine for the default enabled filters to match the default scheduler driver | |
| 17:33:43 | mriedem | dansmith: i didn't see one | |
| 17:33:45 | dansmith | sounded like sdague was going to but I don't know that he did | |
| 17:34:04 | dansmith | mriedem: is it enough to put something in the config help text or do you want to log something? | |
| 17:34:48 | mriedem | i think we'd need both | |
| 17:35:09 | mnaser | cfriesen i'm onto something, with a totally empty hypervisor -- i see this -- AVAILABLE_SIBLINGS: [CoercedSet([1, 17]), CoercedSet([31, 15]), CoercedSet([23, 7]), CoercedSet([13, 29]), CoercedSet([27, 11]), CoercedSet([19, 3]), CoercedSet([9, 25]), CoercedSet([5, 21])] | |
| 17:35:26 | dansmith | mriedem: okay | |
| 17:35:27 | cdent | dansmith: was the resolution on the question of what to do when we get to the end of that set of conditionals where jay had “literally no idea what to do” to do nothing? | |
| 17:35:32 | mriedem | dansmith: plus reno of course | |
| 17:35:32 | mnaser | [1, 17] shouldn't be there, it should just be 17, because vcpu_pin_set is 2-31 | |
| 17:35:43 | mnaser | i dont think that codebase takes vcpu_pin_set into consideration | |
| 17:35:49 | dansmith | cdent: yes | |
| 17:35:55 | cdent | ✔ | |
| 17:36:57 | mriedem | gibi: ok so i guess it is a regression in pike then so i'll mark it pike-rc-potential | |
| 17:37:20 | cfriesen | mnaser: you had changed vcpu_pin_set to 0,16, no? | |
| 17:38:00 | cfriesen | mriedem: okay, I'll respin | |
| 17:38:42 | mnaser | cfriesen i switched back. by setting it to 0,16 => i end up being unable to spin the last instance because all cores + memory are taken on numanode #2 | |
| 17:38:43 | mnaser | s/#2/#1/ | |
| 17:38:49 | mnaser | and numanode 0 has 4gb memory left unused | |
| 17:39:00 | mnaser | by using 0,16 -- i have a single thread in both cores that can use the leftover 4gb | |
| 17:39:28 | mnaser | using 0,16 effectively means 14 threads for 1 numanode and 16 threads for the other. vs 15/15 | |
| 17:40:11 | cfriesen | mnaser: please double-check that you restarted nova-compute or rebooted after changing vcpu_pin_set. I agree that you shouldn't see 1 in the AVAILABLE_SIBLINGS list. | |
| 17:40:13 | mriedem | gdi, people, tox -e fast8 before git review | |
| 17:40:14 | mriedem | https://review.openstack.org/#/c/488510/ | |
| 17:41:03 | mnaser | DEBUG oslo_service.service [req-39a92c29-83e3-4cb0-9f78-29b97d61147d - - - - -] vcpu_pin_set = 2-31 log_opt_values /usr/lib/python2.7/site-packages/oslo_config/cfg.py:2622 in the logs, /proc/cmdline contains isolcpus=2-31 | |
| 17:44:01 | mnaser | http://paste.openstack.org/raw/617965/ this is what is generated from `resources` in claims.py | |
| 17:44:42 | mnaser | i dont even see it in the siblings there, i wonder if it is in the process of the transformation to the DB object | |
| 17:47:15 | cfriesen | mriedem: what do you think of this wording? http://paste.openstack.org/show/617966/ | |
| 17:48:40 | mriedem | "If you have ``scheduler_driver`` set to ``caching_scheduler``, you should probably add these entries back in along with CoreFilter." | |
| 17:48:51 | mriedem | issue the first: scheduler_driver is the old option name | |
| 17:48:57 | mriedem | it's now [scheduler]/driver | |
| 17:49:14 | cfriesen | whoops, was looking at wrong window. :) | |
| 17:49:42 | mriedem | issue the 2nd: let's not tell them what to do if they're not using the filter scheduler, but just point out that non-FilterScheduler drivers, like the CachingScheduler, may want to have these filters enabled | |