| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-09 | |||
| 15:23:23 | mnaser | oh boy | |
| 15:23:37 | mnaser | did i just forget how to math | |
| 15:23:50 | mnaser | and the fact it had 29 cores probably messed up the math | |
| 15:24:06 | danpb | sibling_sets gets populated from available_siblings iiuc | |
| 15:24:15 | danpb | so presumably available_siblings is empty too ? | |
| 15:24:42 | mnaser | danpb: let me put some log.debug's, but i confirmed that siblings_set was empty | |
| 15:24:53 | mnaser | fyi this issue only occurs on the *last* VM to fit on the compute node | |
| 15:26:12 | danpb | well i think you'd need to work backwards through the call stack to figure out why it becomes empty | |
| 15:26:17 | mnaser | danpb: AVAILABLE_SIBLINGS: [CoercedSet([]), CoercedSet([]), CoercedSet([]), CoercedSet([]), CoercedSet([]), CoercedSet([]), CoercedSet([])] | |
| 15:27:19 | mnaser | ok i guess ill have to see why host_cell.free_siblings is being set to that | |
| 15:27:30 | dansmith | danpb: doesn't this come from libvirt? | |
| 15:28:12 | mnaser | the only thing i can imagine which can cause a corner case is the fact that we reserve 2 cores for the OS, so vcpu_pin_set=2-31 .. maybe that's not taken in consideration (guessing) | |
| 15:28:17 | openstackgerrit | Matt Riedemann proposed openstack/nova master: [placement] Add api-ref for usages https://review.openstack.org/480563 | |
| 15:29:51 | danpb | dansmith: libvirt will provide info on the siblings present on the host, but iiuc free_siblings is populated by nova | |
| 15:30:17 | danpb | so if mnaser is saying it works correctly for all VMs until the last one, it sounds like nova is filtering the info from libvirt and ending up with the empty set | |
| 15:30:24 | dansmith | danpb: oh okay, I hadn't traced it very far up the stack because I assumed we must just be getting an empty set of things from libvirt because we had no numa info or something | |
| 15:30:32 | dansmith | danpb: yeah, makes sense | |
| 15:30:35 | sahid | danpb: mnaser danpb i just for information i was working on this but did not find the root cause | |
| 15:30:37 | sahid | https://review.openstack.org/#/c/458848/ | |
| 15:30:42 | mnaser | im checking the values of host_cell and instance_cell | |
| 15:31:15 | mnaser | HOST_CELL: NUMACell(cpu_usage=14,cpuset=set([2,4,6,8,10,12,14,16,18,20,22,24,26,28,30]),id=0,memory=196562,memory_usage=57344,mempages=[NUMAPagesTopology,NUMAPagesTopology],pinned_cpus=set([2,4,6,8,10,12,14,18,20,22,24,26,28,30]),siblings=[set([8,24]),set([2,18]),set([10,26]),set([12,28]),set([6,22]),set([14,30]),set([4,20])]) | |
| 15:31:36 | danpb | how many VMs are you running ? | |
| 15:32:09 | danpb | you've got 7 pairs of siblings there, so if each VM wanted one pair, you'd be able to run 7 VMs | |
| 15:32:42 | mnaser | booted with 120x 1GB hugepages, 30 cores available, trying to boot 15 VMs with 2 cores each + 8gb of memory each | |
| 15:32:43 | mnaser | but im not using isolate, im using prefer which i believe should try to schedule them on the same hyperthread | |
| 15:33:17 | mnaser | my flavor has properties: hw:cpu_policy=dedicated, hw:cpu_thread_policy=prefer, hw:mem_page_size=1048576, hw:numa_nodes=2 | |
| 15:33:58 | mnaser | (also hw:cpu_thread_policy=prefer forgot to put that in) | |
| 15:34:23 | mnaser | i was thinking hw:numa_nodes=2 and hw:cpu_thread_policy=prefer might be the source of the issue because they are a bit the opposite of each other | |
| 15:35:48 | mnaser | getting memory split from 2 numa nodes but also wanting to be on the same hyperthread isnt really possible considering you'll want to be on two physical cpus to access the numa node | |
| 15:36:07 | openstackgerrit | Matt Riedemann proposed openstack/nova master: [placement] Add api-ref for RP usages https://review.openstack.org/450105 | |
| 15:37:50 | danpb | mnaser: how many vcpus per guest ? | |
| 15:38:00 | danpb | (i mean what is vcpu count in the flavour ?) | |
| 15:38:04 | mnaser | danpb: 2 | |
| 15:38:18 | danpb | that flavour config looks flawed then | |
| 15:38:33 | danpb | you've got 2 vcpus, and you've said to give the VM 2 virtual numa nodes | |
| 15:38:55 | danpb | nova will want to place each virtual numa node on a separate host numa node | |
| 15:39:03 | danpb | so each vcpu will need to be on a separate socket | |
| 15:39:09 | openstackgerrit | Matthew Edmonds proposed openstack/nova master: use conf for keystone session creation https://review.openstack.org/485121 | |
| 15:39:16 | danpb | so asking nova to put the vcpus in the same hyperthread sibling is nonsensical | |
| 15:39:29 | danpb | as you can't have siblings if you have separate sockets | |
| 15:39:43 | mnaser | that is what i kinda theorized a bit | |
| 15:39:54 | mnaser | so will have to fallback to 1 numanode default | |
| 15:40:08 | danpb | nova shouldn't be trying to use hypthread siblings at all in ths case | |
| 15:40:43 | mnaser | that was my assumption that what it would do, thats why i put 'prefer' in there because i figured that was the more reasonable choice | |
| 15:40:46 | cfriesen | with hw:cpu_thread_policy=prefer it should just give separate pCPUs. if you had hw:cpu_thread_policy=require I think it'd fail | |
| 15:40:59 | danpb | putting in prefer should be an no-op in this case | |
| 15:41:14 | danpb | as the need to pick separate sockets should have taken priority in nova to satisfy the numa constraint | |
| 15:41:28 | danpb | and if you have thread_policy=require, then nova ought to raise a fatal error | |
| 15:41:59 | cfriesen | hw:cpu_thread_policy=prefer is always a no-op, technically. it's the default behaviour | |
| 15:42:43 | mnaser | right: so in that case, the config seems to be ok (yes, the prefer is noop/will never work) but it should still work | |
| 15:42:52 | mnaser | numastat -m reports 4096 free hugepages in each node (for a total of 8192) | |
| 15:45:23 | cfriesen | too bad sfinucan is on holidays. :) | |
| 15:45:59 | mnaser | the used cores are: 2-15,18-31 | |
| 15:46:12 | mnaser | which leaves 16/17 free which are on two different sockets | |
| 15:46:43 | mnaser | so technically it should schedule on 16,17 which will be able to get 4096MB from each numa socket | |
| 15:47:53 | cfriesen | mnaser: which version of the code is this? | |
| 15:47:58 | mnaser | newton | |
| 15:48:09 | mnaser | python-nova-14.0.7-1.el7.noarch | |
| 15:48:10 | mnaser | more specifically | |
| 15:49:47 | mnaser | sahid i looked at your change however it seems siblings_set being empty is a possiblity :( | |
| 15:50:03 | cfriesen | dansmith: for some reason I thought that pCPUs without a sibling would still be in sibling_set, just as a single-item set | |
| 15:50:46 | efried | sdague https://github.com/openstack/nova/blob/master/nova/cmd/status.py#L198 Am I missing something, or is this untrue? | |
| 15:51:11 | hongbin | oomichi: hi, my team is working on deciding the default api version, a team member asked me to check with you since you did the api version work in nova, i mainly wanted to know how nova maintain the default api version (pick the latest version or a stable version? bump the default version everytime a new version is introduced? etc.) Are you the right person to ask? | |
| 15:51:28 | mnaser | does host_cell get assigned by the scheduler? | |
| 15:52:44 | cfriesen | mnaser: danpb: dansmith: this looks sort of related, but the fix should be in newton: https://bugs.launchpad.net/nova/+bug/1578155 | |
| 15:52:45 | openstack | Launchpad bug 1578155 in OpenStack Compute (nova) newton "'hw:cpu_thread_policy=prefer' misbehaviour" [Medium,Fix committed] - Assigned to Stephen Finucane (stephenfinucane) | |
| 15:52:57 | dansmith | I really don't know anything about this stuff | |
| 15:54:53 | jaypipes | mriedem, dansmith: I will be afk for 2 hours, just FYI. | |
| 15:54:55 | cfriesen | mnaser: so are 16/17 the siblings of 0/1 which are not included in vcpu_pin_set? | |
| 15:55:08 | openstackgerrit | Matt Riedemann proposed openstack/nova master: [placement] Add api-ref for usages https://review.openstack.org/480563 | |
| 15:55:11 | artom | About the fix, at any rate. | |
| 15:55:13 | mnaser | cfriesen nope, 0/1 are the ones which are not included | |
| 15:55:26 | mnaser | so our vcpu_pin_set is 2-31 | |
| 15:55:30 | mnaser | (32 core machine) | |
| 15:55:33 | cfriesen | mnaser: yes, but which are the HT siblings of 0/1? | |
| 15:56:15 | danpb | hmm, yes, it seems numa node 0 has even numbered cpus, and node 1 has odd numbered cpus | |
| 15:56:19 | mnaser | cfriesen sorry, not sure i follow, but here's lscpu output if that helps? http://paste.openstack.org/show/617955/ | |
| 15:56:34 | danpb | so excluding 0-1, kills 1 cpu each host numa node, instead of 2 cpus | |
| 15:56:35 | cfriesen | mnaser: based on the pattern above, the HT siblings of 0/1 would be 16/17....I'm wondering if that's confusing the logic in hardware.py | |
| 15:56:42 | mnaser | but i think that they are | |
| 15:57:19 | cfriesen | mnaswer: running "virsh capabilities" will show sibling info | |
| 15:57:26 | danpb | so that in turn kills a entire sibling pair from each numa node | |
| 15:57:37 | danpb | effectively making 4 cpus unavailable | |
| 15:57:53 | cfriesen | danpb: theoretically it shouldn't kill the sibling pair though...we should have two pCPUs with no siblings | |
| 15:58:03 | mnaser | here is output of virsh caps. http://paste.openstack.org/show/617956/ | |
| 15:58:06 | cfriesen | but maybe the logic can't handle a mix of sibs and no sibs | |
| 15:58:15 | mriedem | someone want to approve this? https://review.openstack.org/#/c/491855/ | |
| 15:58:18 | danpb | yeah that wouldn't be surprising | |
| 15:58:46 | mnaser | so am i better off reserving 2 siblings and maybe that might get rid of the 'confusion' ? | |
| 15:58:53 | cfriesen | mnaswer: seems likely | |
| 15:58:55 | mnaser | so instead of 0,1 .. do 0,16 instead? | |
| 15:59:02 | danpb | yeah | |
| 15:59:53 | mnaser | let me test that theory out | |
| 16:00:03 | openstackgerrit | Ed Leafe proposed openstack/nova master: Handle addition of new nodes/instances in ironic flavor migration https://review.openstack.org/487954 | |
| 16:00:12 | edleafe | dansmith: dtantsur: ^^ updated to remove the node RC change handling | |
| 16:00:37 | dtantsur | awesome! I'm just out of a meeting, will do the ironic part now | |
| 16:00:50 | mnaser | so my vcpu_pin_set should technically be 1-15,17-31 in that case, correct? | |
| 16:00:53 | mnaser | leaving 0,16 for the OS | |
| 16:00:53 | mriedem | gibi: wasn't that also a problem in ocata then? CoreFilter isn't enabled by default | |
| 16:00:59 | cfriesen | mnaswer: either way, please raise a bug for this issue, it'd be nice to handle it properly | |