| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2019-11-20 | |||
| 13:40:55 | sean-k-mooney | so your 8 core vm needs to be spread across 3 cache regions and 2 numa nodes | |
| 13:40:56 | mnaser | sean-k-mooney: it does, when i enable numa node per l3 cache, i end up with 6 threads in a numa node | |
| 13:41:43 | sean-k-mooney | oh ya i guess you could fit it on 2 cache regions and 1 numa node if you are using hyperthreads | |
| 13:42:28 | sean-k-mooney | mnaser: so rather then select core fomr the same numa node if we wanted to optimise in nova we would want to select cores in the same cache region first | |
| 13:42:49 | sean-k-mooney | we can actuly model the cache toploy in libvirt too and expose that to the guest | |
| 13:43:23 | sean-k-mooney | mnaser: it would be interesting to see how a 6 core vm faired against the 8 core vm | |
| 13:43:46 | mnaser | yeah, i read a bit about that. i'm wondering if performance wise, exposing numa node per l3 cache and then doing 2 numa nodes in a VM would result in good overall performance | |
| 13:44:05 | mnaser | you'd have 4 cores each sitting in a numa node and hopefully the guest is smart enough to understand the distances involved | |
| 13:44:49 | sean-k-mooney | yes although again nova does not today fully understand the hierachy | |
| 13:45:03 | sean-k-mooney | e.g. nova could select 2 numa nodes form differnet sockets | |
| 13:45:21 | sean-k-mooney | ideally it woudl select 2 form the same socket to miniums the numa distance | |
| 13:45:35 | sean-k-mooney | that is actully just an optimization of the current behavior | |
| 13:45:43 | sean-k-mooney | rahter then a change | |
| 13:45:49 | mnaser | sean-k-mooney: right but at least the guest will be aware there is possible latency | |
| 13:45:54 | sean-k-mooney | yep | |
| 13:45:55 | mnaser | yep, indeed, its not optimal, but its an improvement | |
| 13:46:46 | sean-k-mooney | so the reason i suggested disabling the multiple numa nodes in zen 1 was our pms insisting vnf dont understand numa | |
| 13:47:08 | sean-k-mooney | after 5+ years of eveyone explaiming what numa is to vnf vendors i dont really buy that | |
| 13:47:57 | mnaser | im wondering which would be better: NPS=4 + no NUMA per L3 + 1 numa node .. vs .. NPS=4 + NUMA per L3 + 2 numa nodes -- im thinking the latter will likely be much faster (or at least the VM will be 'smarter' at understanding the topology.. somewhat) | |
| 13:48:26 | mnaser | it might be unpredictable because the NUMA node might have a distance of 32, or 12, or 10 .. but at least it might know there is distance | |
| 13:48:40 | sean-k-mooney | the later i think would optimize the memory latency | |
| 13:50:21 | mnaser | im gathering some benchmark info | |
| 13:50:29 | mnaser | and then ill try the two numa node + l3 per numa | |
| 13:50:32 | sean-k-mooney | exposing NUMA per L3 should give the best performance it just requires your vms to have multiple numa nodes if they have more then 3-4 cores or >32-64G of ram | |
| 13:50:53 | mnaser | yep, most of the VMs will be 8 core / 8 gb | |
| 13:51:10 | mnaser | also FWIW those aren't dedicated cores | |
| 13:51:39 | sean-k-mooney | sure you just setting hw:numa_nodes=2 but not enabling pinning | |
| 13:51:48 | aarents | hi dansmith, can you confirm that the last update is fine for you ? https://review.opendev.org/#/c/670000 | |
| 13:51:57 | mnaser | yep, os that way im thinknig nova should do 4 cores per NUMA, 4gb in each numa node | |
| 13:52:03 | sean-k-mooney | if you do that you should also set hw:mem_page_size to either large or small | |
| 13:52:27 | sean-k-mooney | mnaser: yes it would | |
| 13:53:00 | sean-k-mooney | the reason for doing hw:mem_page_size by the way is you want to enable novas numa aware memory tracking | |
| 13:53:24 | sean-k-mooney | if your dont want to use hugepages set it to "small" | |
| 13:53:39 | sean-k-mooney | that will use 4k pages | |
| 13:53:46 | mnaser | sean-k-mooney: i think huge pages is soemthing id want, but i think transparent huge pages are enabled right now | |
| 13:54:26 | sean-k-mooney | right but you should still set hw:mem_page_size=small if you have a numa toplogy and are not using explcit hugepages | |
| 13:54:33 | sean-k-mooney | if you dont you can get OOM events | |
| 13:54:45 | mnaser | i think id probably want large anyways, host has 512G memory and plenty of reserved memoryt | |
| 13:55:12 | sean-k-mooney | ya if you dont allow memory oversubsription then there is no reason not to use hugepages | |
| 13:55:32 | mnaser | ya nope no oversubscription | |
| 13:58:52 | mnaser | epyc rome numbers https://www.irccloud.com/pastebin/3bgON0HX/ | |
| 13:59:18 | mnaser | intel (but much older hardware) https://www.irccloud.com/pastebin/6M7B4QLY/ | |
| 13:59:22 | mnaser | sean-k-mooney: ^ fyi, interesting to see latency numbers | |
| 14:01:05 | coreycb | kashyap: do you know if anything has landed for removing cpu flags, similar to cpu_model_extra_flags? | |
| 14:01:28 | kashyap | coreycb: Hi...afraid, no; but I've filed a Blueprint for it: | |
| 14:01:55 | kashyap | coreycb: https://blueprints.launchpad.net/nova/+spec/allow-disabling-cpu-flags | |
| 14:02:30 | coreycb | kashyap: thanks | |
| 14:03:40 | coreycb | kashyap: fyi this bug is why I'm asking: https://bugs.launchpad.net/bugs/1853200 | |
| 14:03:40 | openstack | Launchpad bug 1853200 in libvirt (Ubuntu) "cpu features hle and rtm disabled for security are present in /usr/share/libvirt/cpu_map.xml" [High,Confirmed] - Assigned to Ubuntu Security Team (ubuntu-security) | |
| 14:03:44 | kashyap | coreycb: I know :-( | |
| 14:03:51 | kashyap | And I guessed as much | |
| 14:04:35 | kashyap | coreycb: One (very valid) 'workaround' is that have QEMU add new "named CPU models" to remove the said flags. | |
| 14:05:00 | coreycb | kashyap: ok I'll mention that, thanks | |
| 14:05:43 | kashyap | coreycb: For that, upstream QEMU folks must add them...please file a QEMU "RFE" bug on launchpad for it | |
| 14:06:03 | sean-k-mooney | mnaser: did you enable cluster on die for the intel system out of interest | |
| 14:06:28 | coreycb | kashyap: thanks I'll pass the details on to cpaelzer, he's our qemu maintainer | |
| 14:06:39 | kashyap | (Nod) | |
| 14:06:50 | sean-k-mooney | mnaser: but yes the amd number seam to be much higher | |
| 14:07:17 | sean-k-mooney | mnaser: that is quite surprising to be honest | |
| 14:39:15 | efried | bauzas, gibi, stephenfin: would one of you please have a look at the nova/cyborg spec update https://review.opendev.org/#/c/684151/ and +A if appropriate? | |
| 14:39:35 | bauzas | efried: yeah, I can do | |
| 14:39:40 | efried | thanks bauzas | |
| 14:39:48 | stephenfin | I didn't review the original so I'll defer to others | |
| 14:39:55 | bauzas | I promised to do spec reviews, in particular following the ones we discussed at the PTG | |
| 14:39:58 | efried | After all, melissaml and Viere are +1. | |
| 14:45:43 | efried | dustinc: Did you pick up what I laid down yesterday about compute node conflicts for provider config? http://eavesdrop.openstack.org/irclogs/%23openstack-nova/%23openstack-nova.2019-11-19.log.html#t2019-11-19T22:19:41 | |
| 14:45:59 | aarents | Hi guys I need some review on this https://review.opendev.org/#/c/678016/ gibi already put +2, if you can have a look ? | |
| 14:46:55 | bauzas | aarents: /me clicks | |
| 14:47:24 | efried | dansmith: TL;DR: since, ironic notwithstanding, it would be insane for one compute node's $name to be the same as another's $uuid, we're just going to punt if there's a conflict across the whole $name+$uuid space. With that in play, it makes sense to treat $COMPUTE_NODE as a "default", which can be overridden by a specific $name/$uuid. And all that means we can detect conflicts immediately (on startup) and fail (to start up) if o | |
| 14:48:40 | dansmith | buffer overflow | |
| 14:50:10 | aarents | thks | |
| 14:51:21 | sean-k-mooney | efried: ill try and review the cyborg spec shortly | |
| 14:51:39 | efried | sean-k-mooney: thanks. It's just an update to the existing spec, nothing earth-shaking. | |
| 14:52:02 | dansmith | efried: your message was too long and was cut off after "fail to (start up)" | |
| 14:52:04 | sean-k-mooney | ya looks short | |
| 14:52:28 | efried | dansmith: oh, weird. "...fail (to start up) if one is detected." | |
| 14:52:33 | efried | that was all. | |
| 14:52:51 | efried | my buffer says I had another 25c | |
| 14:52:56 | efried | clearly I need to work on my TL;DRing. | |
| 14:53:11 | dansmith | efried: okay, so I think I was (and likely still am) missing some context yesterday and I was hurrying, | |
| 14:53:34 | dansmith | but the problem is that ironic nodes specified there may not exist as providers on startup, is that it? | |
| 14:53:53 | efried | that's pretty much the only scenario that makes us have to deal with this, yeah. | |
| 14:54:23 | dansmith | so we wanted to make that hard-fail if you provide something that doesn't exist there yeah? | |
| 14:54:28 | efried | no | |
| 14:54:48 | efried | generically the "problem" is that theoretically you can specify a config by $name and another by $uuid, but those are for the same node, but we wouldn't know that until that node "appeared". | |
| 14:55:04 | efried | In practice this shouldn't be possible because for ironic, $name == $uuid. | |
| 14:56:03 | dansmith | you mean if you had two entries in the list, one by name and one by uuid and they were in fact the same provider? | |
| 14:56:08 | efried | right | |
| 14:56:44 | dansmith | that seems like a really tiny detail to be concerned about.. did this come up in some testing or something? | |
| 14:56:53 | efried | more code inspection | |
| 14:57:06 | dansmith | remind me why we have the by-name option anyway? | |
| 14:57:10 | efried | ikr | |
| 14:57:27 | dansmith | uuid or the "self" option should really be all we need I would think | |
| 14:57:30 | efried | It was because you don't necessarily know the UUID yet in a green field. | |
| 14:57:44 | efried | but you want your tripleo to be able to lay down the config a priori. | |
| 14:57:46 | dansmith | in which case, non-ironic? in that case you use $COMPUTE_NODE yeah? | |
| 14:58:12 | efried | Yeah, I would think so. gibi, can we convince you we don't need identification by name? | |
| 14:58:23 | efried | (I think it was gibi who talked us into it) | |
| 14:58:26 | dansmith | either way, | |
| 14:58:46 | dansmith | I'm not sure we really need to care that much if someone puts a thing in there twice | |