Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-09
16:06:11 dansmith mriedem: okay
16:06:32 mriedem gibi: the fix in ocata would have to be different probably since this code all got refactored in pike
16:06:42 mriedem gibi: but just wanted to make sure it's not an rc1 regression blocker thing
16:08:18 dansmith mriedem: I actually think we probably should enable the cache.. I thought it was already done on the computes, but it's not. The other use on compute is to calculate the rpc pin, which is cached until restart anyway
16:09:14 mnaser danpb cfriesen -- i'm now getting "Not enough available CPUs to schedule instance. Oversubscription is not possible with pinned instances. Required: 1, actual: 0"
16:09:28 mnaser HOST_CELL: NUMACell(cpu_usage=14,cpuset=set([2,4,6,8,10,12,14,18,20,22,24,26,28,30]),id=0,memory=196562,memory_usage=57344,mempages=[NUMAPagesTopology,NUMAPagesTopology],pinned_cpus=set([2,4,6,8,10,12,14,18,20,22,24,26,28,30]),siblings=[set([8,24]),set([2,18]),set([10,26]),set([12,28]),set([6,22]),set([14,30]),set([4,20])])
16:11:34 danpb where's your other numa cell
16:11:44 danpb that just shows the first cell
16:12:16 mnaser i dont know why its not listed. i added LOG.debug("HOST_CELL: %s" % host_cell) to _numa_fit_instance_cell_with_pinning
16:12:40 cfriesen mnaser: I think I know what's going on...you now have fewer pCPUs in one node than the other, and you're asking for 2-node guests
16:12:42 mnaser let me get you
16:12:43 mnaser both host cell output
16:12:57 mnaser oh
16:13:01 mnaser i think you're right
16:13:11 danpb cfriesen: yep makes sense
16:13:47 mnaser let me switch things back to how they were and get the output of both host cells
16:14:02 danpb mnaser: why do you want the guests to have multiple virtual numa cells ?
16:14:46 cfriesen mriedem: for a reno for https://review.openstack.org/#/c/491854/ would we want to describe the removal of the two default filters in "features", "upgrade", or "other"?
16:14:52 danpb its generally not something you'd do unless guest memory exceeds the amount available in a single host, or need to consume say, PCI devices from separate nods at the same time
16:15:08 mnaser danpb: still experimenting but the idea was more efficent use of hardware. i have 120x 1gb large pages which means that i'll end up with 60gb on each numanode, if i put 8gb sized instances only, i'll end up with 4096mb in each numa node that's unused
16:15:13 cfriesen danpb: or you want increased memory bandwidth
16:15:54 danpb cfriesen: that's only increased if your guest workload avoids cross-node traffic
16:16:01 cfriesen danpb: agreed
16:16:01 mnaser by doing this, it splits 4096 into each numanode and i can fill the freepages .. but it's not something that's set in stone fully
16:16:11 danpb cfriesen: otherwise you'd actually decrease throughput
16:16:22 cfriesen mnaser: the downside is that your guests need to be able to avoid cross-numa traffic, which makes guest coding trickier
16:16:24 danpb by having the guest all contend on the cross-node memory bus
16:16:48 cfriesen danpb: yes, it'd be a special-case
16:17:00 mnaser yeah i was thinking the linux kernel already does an ok job handling multiple numa nodes
16:17:01 danpb IOW unless your guest app is intelligent you want to avoid multiple numa nods
16:17:10 mnaser i see
16:17:28 cfriesen mnaser: the kernel does, but not all userspace code does.
16:17:38 mriedem dansmith: ack - gonna be afk for about an hour
16:17:40 mnaser also the only other tradeoff that comes with this is that the threads are not shared which is not ideal
16:17:46 mnaser so that was something i didnt like about doing this
16:17:49 mriedem can't be worst dad of the summer 2 days in a row
16:17:57 cfriesen mnaser: so if you've got a single large app that wants most of that 8GB...
16:18:16 dansmith mriedem: okay I'll push this up with those changes and I'll be gone before jenkins gets to it
16:18:23 mnaser for completition sake btw, this is both host cells
16:18:24 dansmith mriedem: jaypipes so baton back to you at that point
16:18:45 mnaser http://paste.openstack.org/show/617957/
16:19:26 mnaser there are 15 sets there
16:19:37 mnaser let me see what it was when we had 0-1 only
16:20:17 cdent dansmith: couple questions about expected discover_hosts behavior: is it supposed to be idempotent (run it again and again, it’s okay)? It is supposed to cope if two different processes run it at the same time?
16:20:24 cfriesen mnaser: but only 14 pairs of pCPUs from different numa nodes
16:20:43 mnaser yeah, thats why it failed now, but in the original case
16:20:44 mnaser strangely enough
16:20:51 mnaser i only see one output of numacell, not two
16:20:51 dansmith cdent: yeah, it's expected to just run it over and over again, from cron even
16:21:15 dansmith cdent: you might get a failure if you run multiples and one loses the race to insert the record, but other than that it should be fine to run multiple threads of it
16:21:41 cdent dansmith: thanks that’s what I was hoping/expecting but wanted to confirm
16:23:51 mnaser im going to try setting hw:numa_nodes=1 and packing the server and seeing what happens (i should be able to get 14 at least)
16:25:47 gibi mriedem: ocata is not affected by the bug: https://github.com/openstack/nova/blob/stable/ocata/nova/scheduler/filter_scheduler.py#L188
16:26:19 gibi mriedem: I mean bug 1708637
16:26:21 openstack bug 1708637 in OpenStack Compute (nova) "nova does not properly claim resources when server resized to a too big flavor" [High,In progress] https://launchpad.net/bugs/1708637 - Assigned to Balazs Gibizer (balazs-gibizer)
16:26:33 cfriesen mnaser: do you actually need 8GB? if you don't actually need it all, you could drop to 2MB hugepages and divide up the memory evenly with less waste. For most things 1GB pages don't give that big of a boost.
16:27:35 mnaser cfriesen have you had experience with it? i just figured that if i can have 1gb pages, it would be better than 2mb pages but there isn't much substance to it other than 'it seems right'
16:27:42 mnaser dropping to 2mb would obviously make life much easier
16:27:54 cdent mriedem: yeah, we’ve been pretty inconsistent about which 400s are document in the placement-api-ref. anything common like “yo, not here” and “hey, you violated schema” has frequently been dropped. I’ve not been too strict on my reviews of that stuff except where a 4xx has some particular weird sense
16:28:39 cdent the current tooling doesn’t really have as much support for handling error responses as we might want, that’s probably something we could and should address later. I think once the whole thing is documented we’ll be able to tune it as a whole, better
16:29:14 mnaser i just tried to create 14x 2vcpu/8gb memory and that went ok, then 1x 1vcpu/4gb memory and that was okay, so now i have 4gb free large pages.. trying to create 1vcpu/4gb - Host does not support requested memory pagesize. Requested: 1048576 kB
16:30:42 mnaser looks like it was assigned core 21, which is on node1 .. but my free 4096 is on node0
16:30:53 mnaser i guess this happened because im still using the same thread pair, probably wouldnt happen again if i reserve 0-1 for the os
16:31:54 cfriesen mnaser: we've done some testing. there are some cases where it makes a noticeable difference, but most of the time the difference is not worth the wasted memory. it depends on the guest
16:33:02 mnaser cfriesen all the guests are going to be multiples of 1gb in memory so 120gb in large pages in 1gb or in 2mb .. wouldn't make much difference, no? it looks like the # of freepages is split across both numa nodes
16:34:59 cfriesen mnaser: if it must be exact multiples of 1GB, then you may as well use 1GB hugepages.
16:35:43 mnaser cfriesen yeah they're all multiples of 1gb.. anyways ill try going back to 0-1 and seeing if i can fully populate the server
16:35:55 mnaser and then after that ill file a bug regarding that sibling_set issue
16:44:33 cdent mriedem: is this still relevant or has other stuff killed it: https://review.openstack.org/#/c/488187/
17:13:05 mnaser reserving 0-1 allows me to start 14 2vcpu/8gb, but refuses to let me start anymore with that same sibling_sets bug
17:14:19 openstackgerrit Chris Friesen proposed openstack/nova master: Remove ram/disk sched filters from default list https://review.openstack.org/491854
17:17:16 openstackgerrit Chris Dent proposed openstack/nova master: replace chance with filter scheduler in func tests https://review.openstack.org/491529
17:18:21 mnaser cfriesen: doing more debugging, i believe i'm onto something -- HOST_TOPOLOGY: NUMATopology(cells=[NUMACell(UNKNOWN),NUMACell(1)])
17:18:42 mnaser for some reason the first numacell is unknown..? i'll have to check that
17:23:23 melwitt dansmith: +1 to skipping cells meeting
17:29:17 mriedem cdent: not sure, but not something we need to care about for rc1
17:29:52 cdent mriedem: ‘k, will circle back round to that after we get the main things settled
17:31:20 openstackgerrit Dan Smith proposed openstack/nova master: Resource tracker compatibility with Ocata and Pike https://review.openstack.org/491012
17:31:30 dansmith mriedem: jaypipes ^
17:33:00 mriedem cfriesen: commented in your change
17:33:16 mriedem cfriesen: if you're using the caching scheduler, you still rely on those filters since the caching scheduler doesn't use placement
17:33:37 dansmith mriedem: did we ever push up that deprecation for those schedulers?
17:33:38 mriedem but the caching scheduler isn't the default filter driver, the filter_scheduler is, so i think it's fine for the default enabled filters to match the default scheduler driver
17:33:43 mriedem dansmith: i didn't see one
17:33:45 dansmith sounded like sdague was going to but I don't know that he did
17:34:04 dansmith mriedem: is it enough to put something in the config help text or do you want to log something?
17:34:48 mriedem i think we'd need both
17:35:09 mnaser cfriesen i'm onto something, with a totally empty hypervisor -- i see this -- AVAILABLE_SIBLINGS: [CoercedSet([1, 17]), CoercedSet([31, 15]), CoercedSet([23, 7]), CoercedSet([13, 29]), CoercedSet([27, 11]), CoercedSet([19, 3]), CoercedSet([9, 25]), CoercedSet([5, 21])]
17:35:26 dansmith mriedem: okay
17:35:27 cdent dansmith: was the resolution on the question of what to do when we get to the end of that set of conditionals where jay had “literally no idea what to do” to do nothing?
17:35:32 mriedem dansmith: plus reno of course
17:35:32 mnaser [1, 17] shouldn't be there, it should just be 17, because vcpu_pin_set is 2-31
17:35:43 mnaser i dont think that codebase takes vcpu_pin_set into consideration
17:35:49 dansmith cdent: yes
17:35:55 cdent ✔
17:36:57 mriedem gibi: ok so i guess it is a regression in pike then so i'll mark it pike-rc-potential
17:37:20 cfriesen mnaser: you had changed vcpu_pin_set to 0,16, no?
17:38:00 cfriesen mriedem: okay, I'll respin
17:38:42 mnaser cfriesen i switched back. by setting it to 0,16 => i end up being unable to spin the last instance because all cores + memory are taken on numanode #2
17:38:43 mnaser s/#2/#1/

Earlier   Later