Earlier  
Posted Nick Remark
#openstack-nova - 2018-03-19
16:20:35 efried As I commented on https://review.openstack.org/#/c/534339/ the 1GB-from-each-of-two-children-to-make-one-2GB-instance thing shouldn't be a problem.
16:20:48 efried BUT the "get my CPU from one RP and my disk from another" would be.
16:21:14 dansmith efried: currently the former is a problem with the vmware driver AFAIK
16:22:03 dansmith cdent: implementing the structurally separated resources in a cluster as nested providers of the root would make progress towards fixing that ^ and then makes the tenant isolation thing a smaller delta we can probably have a more reasonable discussion over
16:22:11 efried To mitigate in UPT-land, you would have to do subtree lassoing like we talked about needing for NUMA; or you would have to model the cluster members as NOT being in the same tree.
16:22:17 cdent dansmith: that's not what the spec does
16:22:24 cdent resource pools are not physical
16:22:39 dansmith efried: that's what I'm saying would be an improvement over what it does today
16:22:45 efried which one dansmith?
16:22:57 cdent if you have a problem with the cluster presented agglomerated resources as a single thing, then you'd have a problem with resource pools too
16:23:01 dansmith efried: representing cluster members as children under the root
16:23:12 cdent max_unit, on either the cluster or the resource pool "fixes" that
16:23:30 efried dansmith: Only if we have the subtree-lassoing technology. I haven't caught up on my specs yet - did someone propose that yet?
16:23:33 dansmith cdent: max_unit does not prevent nova from thinking it can get cpu and memory from two different cluster members
16:23:39 efried Agree ^^
16:24:02 cdent dansmith: yes, and?
16:24:11 dansmith efried: right, I'm assuming the lassoing thing as well as changing this vmware RP exposure.. the two together would be required
16:24:55 dansmith cdent: so max_unit has not "fixed" the vmware driver reporting a whole cluster as a single RP/node
16:24:56 efried In that case, yes, I agree we can do this with nested.
16:25:02 cdent the only way we get what you seem to want is for every esxi host to represent all its resources, which breaks the DRS, unless the virt driver can write allocations
16:25:26 dansmith cdent: yeah, I think we've asserted that DRS under nova is broken, for that reason exactly
16:25:34 dansmith broken fundamentally I mean
16:25:59 efried cdent: I thought we talked about the fact that you want to hide the whole cluster-ness anyway?
16:26:20 efried Represent the whole thing as a single RP, and then vmware virt would do the individual node business under the covers.
16:26:32 cdent efried: yes, dan's saying that's broken
16:26:34 dansmith efried: that's what it does today, and it's breakable
16:26:58 dansmith efried: because nova will not know that one node is out of CPU but has some memory available, where another node has the opposite.. nova will think it can schedule to that, but it can't
16:27:07 gjayavelu @cdent @dansmith @efried I'm catching up on the discussions about the resource pools spec
16:27:45 efried dansmith: Okay, I agree with that. But I think the virt can probably mitigate
16:27:53 dansmith efried: it can't
16:27:55 efried ...by cleverly spoofing the inventories.
16:28:05 dansmith oh, it could work around it that way sure
16:28:13 dansmith that's fairly wasteful though :)
16:28:17 efried Basically by representing the least common denominator of the pool, as the inventory.
16:28:24 efried Is it?
16:28:30 dansmith sure
16:28:33 efried Because you wouldn't be able to schedule more than that anyway, wouldja?
16:28:59 dansmith well, maybe you're right
16:29:06 cdent yeah, you either fail to get allocation candidates, or you fail later
16:29:07 dansmith well, no,
16:29:13 cdent better to fail to get allocation candidates
16:29:43 dansmith you could have a node with small memory, small cpu, lots of disk, and another node with zero disk but lots of other resources.. you have to report yourself as fully committed at that point right?
16:29:47 efried Okay, I see, if you had one node with 1VCPU/1024MEM and one with 2VCPU/512MEM, you would wind up representing 1/512 even though you could technically schedule 1/1024 or 2/512.
16:29:49 dansmith cdent: also that
16:30:00 efried dansmith: Yeah, I think we're saying the same thing.
16:30:04 dansmith yar
16:30:23 cdent at the moment that is an accepted shortcoming
16:30:35 cdent mostly because the common case is for everything in the cluster to be the same
16:30:58 efried Yeah, was gonna say, homogeneity would make most of the issue moot.
16:31:02 dansmith right but your flavors have to basically fit together perfectly like puzzle pieces to avoid too much fragmentation
16:31:15 cdent ENOPARSE
16:31:18 openstackgerrit sahid proposed openstack/nova master: libvirt: slow live-migration to ensure network is ready https://review.openstack.org/497457
16:31:35 efried I think it's a pretty solid 80/20 that's worth going forward with.
16:32:32 efried Hypothetically the DRS will also reshuffle things to maintain/restore homogeneity, nah? That's kind of its job?
16:32:54 cdent yes
16:32:58 cdent well
16:33:09 cdent it will shuffle things. but it can't move hardware
16:33:21 dansmith that doesn't fix anything if your flavors are different sizes
16:33:21 efried understood.
16:33:27 efried It can
16:33:46 efried because it can consolidate VMs in ways that maximize utilization of a node.
16:33:50 efried If it's clever enough.
16:34:25 dansmith it may be able to if things fit right, yes
16:34:26 cdent even if its not clever enough, if vmware is being wasteful, that's vmware's problem, right?
16:34:27 efried But anyway, IMO it's acceptable to state that we know this solution is not 100% perfect, and move forward with it, because it will work *mostly* well, *most* of the time.
16:34:45 openstackgerrit Merged openstack/nova master: Revert "Refine waiting for vif plug events during _hard_reboot" https://review.openstack.org/553035
16:46:05 melwitt cool, thanks gibi
16:48:01 melwitt gibi, bauzas: I'd like to bring to your attention the draft for runways during the rocky cycle, if you have any feedback about it https://etherpad.openstack.org/p/nova-runways-rocky I'd like to kick of the process later this week so we can try it out and adjust it as we go
16:53:07 gibi melwitt: I opened that etherpad and I will try to check it tomorrow
16:53:17 gibi melwitt: seem like a pretty comprehensive doc
16:54:41 melwitt gibi: cool, thanks. yeah, there's been a lot of feedback already last week, so np if there's nothing else you'd like to add or ask
16:55:19 gibi melwitt: I will try to read it anyhow :)
16:55:36 melwitt thanks. I'll ping bauzas again tomorrow as it looks like he's on PTO today
16:57:02 openstackgerrit Eric Berglund proposed openstack/nova master: PowerVM Driver: vSCSI volume driver https://review.openstack.org/526094
16:58:35 gibi mriedem_away, alex_xu_ , mlavalle: thanks for the awesome review feedback on the bandwidth spec. I tried to answer the questions inline https://review.openstack.org/#/c/502306
16:59:24 mlavalle gibi: will take a look again soon. Thanks!
17:00:26 gibi mlavalle: thanks
17:03:09 openstackgerrit Ken'ichi Ohmichi proposed openstack/nova master: Remove version/date from CLI documentation https://review.openstack.org/553903
17:10:28 jaypipes sorry y'all. off phone call now. reading back..
17:13:52 tssurya dansmith: I had a question about the reset cache implementation on the scheduler manager as a part of the cell disable spec, is it a good time now ? If you are busy I can come back later
17:14:05 dansmith tssurya: sure
17:14:31 tssurya So I was implementing the reset cache option on the scheduler manager , however when using the scheduler client from nova manage to call this reset on the cache in the manager, doesn’t this become an upcall which we don’t support ?
17:15:03 tssurya or am I not supposed to call this using the client from nova-manage ?
17:15:12 dansmith tssurya: I was saying this should be done in a reset() handler on scheduler manager, which gets triggered on SIGHUP
17:15:30 dansmith tssurya: correct, not from nova-manage or anything else, I thought I commented to that effect on the spec
17:15:49 tssurya yes you did, I guess I didn't understand the SIGHUP very well
17:15:53 dansmith tssurya: this is the example of a similar thing we have: https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L543-L546
17:16:17 dansmith tssurya: so you can SIGHUP compute manager now to get it to clear the service version cache and rpc pin, which is the same sort of activity you're doing
17:17:21 tssurya right,
17:17:47 tssurya I need to "SIGHUP" the scheduler manager basically
17:20:36 dansmith tssurya: yeah, we have a couple other signal handlers for things like guru meditation, etc
17:33:40 lennyb Hi, my nova instance got stuck in 'deleting' when I try to delete instance that failed to be deployed via ironic n-cell-child.service.log
17:33:41 lennyb http://paste.openstack.org/show/704774/
17:41:16 openstackgerrit Eric Fried proposed openstack/nova-specs master: Mention (no) granular support for image traits https://review.openstack.org/554305
17:42:33 efried jaypipes, mriedem: ^
17:46:23 oomichi melwitt: cdent: hi, do we have an approval at the PTG for https://blueprints.launchpad.net/nova/+spec/placement-extract ?
17:46:27 jaypipes dansmith, mriedem: would you mind looking at https://review.openstack.org/#/c/534339/ please and letting me know if you agree with efried? I'm thinking I agree with him and if so, will abandon that and the following patch to it.
17:46:54 cdent oomichi: the decision at ptg was to make progress in rocky, but not complete it, and not as a priority
17:47:57 oomichi cdent: I see, thanks. I just wanted to see the bp as approved
17:48:54 cdent oomichi: I didn't actually create the bp until late last week, because efried suggested it would be useful. We hadn't really declared "let's have one" but it does seem like a good idea.

Earlier   Later