Earlier  
Posted Nick Remark
#openstack-nova - 2018-06-06
15:05:37 dansmith bauzas: right that's the point
15:06:11 bauzas dansmith: but https://review.openstack.org/#/c/572195/4/nova/conf/scheduler.py is asking for a default number that's huge
15:06:33 dansmith bauzas: did you read the justification for that number in the commit message?
15:06:45 bauzas sec, I maybe missed it
15:06:53 bauzas if so, my bad
15:10:06 dansmith if it's wrong, please do call it out, but it _was_ an intentional plan :)
15:10:29 bauzas dansmith: /me looking at the related bug to understand the problem
15:12:23 dansmith the first paragraph pretty much sums it up, but the related bug is also good context
15:14:45 bauzas dansmith: humpf, I wasn't knowing https://github.com/openstack/nova/commit/f93f675a :)
15:15:07 dansmith really? okay :)
15:15:15 bauzas yeah, really
15:15:25 dansmith you were in boston, no?
15:15:44 bauzas ok, so now I understand why we want to shuffle computes per number of failed builds
15:16:06 bauzas dansmith: yup, sorry but I don't remember that discussion :(
15:16:08 openstackgerrit Curt Moore proposed openstack/nova-specs master: Add spec for downloading images via RBD https://review.openstack.org/572805
15:16:11 dansmith okay
15:16:30 bauzas at least, I now understand the problem
15:21:02 bauzas dansmith: in general, we said a couple of times to operators to randomize more the computes by using host_subset_size or shuffle_best_same_weighed_hosts
15:21:20 bauzas dansmith: so that a new instance was asking for a separate compute
15:21:30 dansmith yeah, but people that want strict packing might not want that
15:21:38 bauzas dansmith: agreed
15:21:45 jaypipes efried: done.
15:21:46 bauzas asking to pack is a problem
15:21:54 efried jaypipes: thanks
15:22:13 bauzas dansmith: what I was more thinking was maybe to deprecate one of the two options I mentioned
15:22:23 bauzas dansmith: and recommend using your own weigher
15:22:47 bauzas because once we merge your weigher, we'll get 3 opts for quite the same concern
15:23:06 dansmith um, why?
15:23:10 dansmith I don't think those overlap
15:23:40 bauzas dansmith: I think shuffle_best_same_weighed_hosts overlaps
15:24:07 dansmith with my weigher?
15:24:20 bauzas dansmith: it's two different implementations
15:24:27 dansmith I don't get that
15:24:38 bauzas but at the end, we want to make sure we pack correctly
15:25:10 dansmith without this weigher, hosts that have been failing will be considered at the same score as those that haven't.. shuffle just changes the likelihood you'll land on one of those, modulo the subset size
15:25:26 dansmith with my weigher they're likely not in consideration, moving your chance closer to 100%
15:26:00 bauzas right, that's my point
15:26:17 bauzas asking to shuffle is just because you want to make sure you go elsewhere
15:26:35 dansmith so, without this weigher,
15:26:44 dansmith lets say your subset size is 3
15:26:48 dansmith and shuffle is turned on
15:27:09 dansmith if you have a few hosts that are empty, they score high. if one of those is misconfigured,
15:27:18 dansmith your chance of booting is 66%
15:27:41 dansmith with this weigher, you'd get two empty hosts, plus one that has stuff on it in the subset
15:27:44 bauzas right, I'm saying your weigher improves the changes
15:27:47 bauzas chances*
15:28:06 bauzas that's why I don't see the need for shuffle_best_same_weighed_hosts
15:28:25 bauzas I'm asking to deprecate shuffle_best_same_weighed_hosts
15:28:49 bauzas because this option just does a random shuffle
15:28:52 dansmith but, in the absence of failing hosts, shuffle still helps to balance concurrent boot traffic right?
15:29:04 dansmith I don't see shuffle as being used for avoiding faults at all
15:32:07 bauzas maybe my concern is just that I dislike this shuffle option :)
15:32:46 bauzas if you want to spread on a specific metric, just ask for a weigher using that metric
15:32:57 bauzas instead of just hoping you'll get spread for free
15:32:59 bauzas anyway
15:34:45 dansmith shuffle is purely to avoid super strict packing right?
15:34:49 dansmith if you have big computes and small instances,
15:35:04 dansmith you might build 100 instances in the same exact compute before it's completely full and you move onto the next bucket in the line
15:35:21 dansmith shuffle seems good for just saying "we generally want packing, but not suuuuuper strict"
15:38:37 bauzas dansmith: "life, uh, always finds a way". Replace "life" by "weigher" and you'll get my mind here :)
15:39:00 bauzas anyway, back to review
15:43:10 openstackgerrit Mathieu Gagné proposed openstack/nova-specs master: Update multiple fixed-IPs support with services field https://review.openstack.org/572814
15:46:52 openstackgerrit Jay Pipes proposed openstack/nova-specs master: Standardize CPU resource tracking https://review.openstack.org/555081
15:55:28 mriedem dansmith: for the follow up https://review.openstack.org/#/c/572814/
15:56:40 dansmith mriedem: don't we have to wait two years first?
15:56:47 dansmith like fine aged cheddar
15:57:04 mriedem artom linked a bz i don't have access to
16:00:00 dansmith mriedem: I looked and I don't think it has any useful context in it
16:00:54 mriedem ok
16:01:07 mriedem i'll proceed with todo cleanupification
16:02:20 openstackgerrit Merged openstack/nova master: Fix doc nit https://review.openstack.org/572682
16:02:44 openstackgerrit Curt Moore proposed openstack/nova-specs master: Add spec for downloading images via RBD https://review.openstack.org/572805
16:03:47 mriedem johnthetubaguy: jaypipes: bauzas: the vmware live migration spec needs another +2 https://review.openstack.org/#/c/299207/
16:04:03 bauzas mriedem: looking
16:05:41 openstackgerrit Curt Moore proposed openstack/nova-specs master: Add spec for downloading images via RBD https://review.openstack.org/572805
16:07:09 openstackgerrit Merged openstack/nova-specs master: Update multiple fixed-IPs support with services field https://review.openstack.org/572814
16:08:04 bauzas mriedem: in pre_live_migration() on the dest compute, we don't manage exceptions raised by the virt driver, right?
16:08:21 bauzas mriedem: I think we only do this kind of pre-flight check in the conductor layer
16:10:13 efried stephenfin: You around? I need some undumbification about tunnel providers.
16:10:20 stephenfin efried: shoot
16:10:30 efried stephenfin: https://review.openstack.org/#/c/564440/3/nova/conf/neutron.py
16:10:38 mriedem bauzas: pre_live_migration is an rpc call from the source compute
16:10:48 mriedem so yes if that fails, the source sets the migration as an error and stops
16:11:17 mriedem bauzas: https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L6077
16:11:32 efried stephenfin: What I thought I saw there is that you can only have your tunnel provider (singular) associated with one NUMA node. Is that right?
16:11:50 stephenfin efried: Nope
16:12:03 stephenfin I'm using the term tunnel provider but what I really mean is tunnel endpoint
16:12:17 bauzas mriedem: ok, so that will work, because I wasn't seeing any exception coming from the virt layer handled by the dest host
16:12:26 stephenfin namely, the logical NIC where all tunneled traffic ingresses and egresses from the compute node
16:12:42 bauzas mriedem: the spec is a bit missing of details, but okay
16:12:44 stephenfin logical because it can actually be two or more physical NICs bonded together
16:13:02 efried stephenfin: Ah. And those multiple NICs may live on multiple NUMA nodes.
16:13:07 openstackgerrit Murali Annamneni proposed openstack/nova master: Enables MySQL Cluster Support for Nova https://review.openstack.org/446643
16:13:13 stephenfin you're only allowed one tunnel endpoint in neutron. Every tunneled network must use that endpoint
16:13:18 stephenfin efried: Correct
16:13:34 stephenfin Ditto for physnet-type networks (e.g. layer 2)
16:14:08 stephenfin You could have bonded NICs spread across the various host NUMA nodes
16:14:25 efried k, I get that.
16:14:48 efried Thanks stephenfin.
16:15:10 stephenfin no problemo

Earlier   Later