| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-06-06 | |||
| 15:16:08 | openstackgerrit | Curt Moore proposed openstack/nova-specs master: Add spec for downloading images via RBD https://review.openstack.org/572805 | |
| 15:16:11 | dansmith | okay | |
| 15:16:30 | bauzas | at least, I now understand the problem | |
| 15:21:02 | bauzas | dansmith: in general, we said a couple of times to operators to randomize more the computes by using host_subset_size or shuffle_best_same_weighed_hosts | |
| 15:21:20 | bauzas | dansmith: so that a new instance was asking for a separate compute | |
| 15:21:30 | dansmith | yeah, but people that want strict packing might not want that | |
| 15:21:38 | bauzas | dansmith: agreed | |
| 15:21:45 | jaypipes | efried: done. | |
| 15:21:46 | bauzas | asking to pack is a problem | |
| 15:21:54 | efried | jaypipes: thanks | |
| 15:22:13 | bauzas | dansmith: what I was more thinking was maybe to deprecate one of the two options I mentioned | |
| 15:22:23 | bauzas | dansmith: and recommend using your own weigher | |
| 15:22:47 | bauzas | because once we merge your weigher, we'll get 3 opts for quite the same concern | |
| 15:23:06 | dansmith | um, why? | |
| 15:23:10 | dansmith | I don't think those overlap | |
| 15:23:40 | bauzas | dansmith: I think shuffle_best_same_weighed_hosts overlaps | |
| 15:24:07 | dansmith | with my weigher? | |
| 15:24:20 | bauzas | dansmith: it's two different implementations | |
| 15:24:27 | dansmith | I don't get that | |
| 15:24:38 | bauzas | but at the end, we want to make sure we pack correctly | |
| 15:25:10 | dansmith | without this weigher, hosts that have been failing will be considered at the same score as those that haven't.. shuffle just changes the likelihood you'll land on one of those, modulo the subset size | |
| 15:25:26 | dansmith | with my weigher they're likely not in consideration, moving your chance closer to 100% | |
| 15:26:00 | bauzas | right, that's my point | |
| 15:26:17 | bauzas | asking to shuffle is just because you want to make sure you go elsewhere | |
| 15:26:35 | dansmith | so, without this weigher, | |
| 15:26:44 | dansmith | lets say your subset size is 3 | |
| 15:26:48 | dansmith | and shuffle is turned on | |
| 15:27:09 | dansmith | if you have a few hosts that are empty, they score high. if one of those is misconfigured, | |
| 15:27:18 | dansmith | your chance of booting is 66% | |
| 15:27:41 | dansmith | with this weigher, you'd get two empty hosts, plus one that has stuff on it in the subset | |
| 15:27:44 | bauzas | right, I'm saying your weigher improves the changes | |
| 15:27:47 | bauzas | chances* | |
| 15:28:06 | bauzas | that's why I don't see the need for shuffle_best_same_weighed_hosts | |
| 15:28:25 | bauzas | I'm asking to deprecate shuffle_best_same_weighed_hosts | |
| 15:28:49 | bauzas | because this option just does a random shuffle | |
| 15:28:52 | dansmith | but, in the absence of failing hosts, shuffle still helps to balance concurrent boot traffic right? | |
| 15:29:04 | dansmith | I don't see shuffle as being used for avoiding faults at all | |
| 15:32:07 | bauzas | maybe my concern is just that I dislike this shuffle option :) | |
| 15:32:46 | bauzas | if you want to spread on a specific metric, just ask for a weigher using that metric | |
| 15:32:57 | bauzas | instead of just hoping you'll get spread for free | |
| 15:32:59 | bauzas | anyway | |
| 15:34:45 | dansmith | shuffle is purely to avoid super strict packing right? | |
| 15:34:49 | dansmith | if you have big computes and small instances, | |
| 15:35:04 | dansmith | you might build 100 instances in the same exact compute before it's completely full and you move onto the next bucket in the line | |
| 15:35:21 | dansmith | shuffle seems good for just saying "we generally want packing, but not suuuuuper strict" | |
| 15:38:37 | bauzas | dansmith: "life, uh, always finds a way". Replace "life" by "weigher" and you'll get my mind here :) | |
| 15:39:00 | bauzas | anyway, back to review | |
| 15:43:10 | openstackgerrit | Mathieu Gagné proposed openstack/nova-specs master: Update multiple fixed-IPs support with services field https://review.openstack.org/572814 | |
| 15:46:52 | openstackgerrit | Jay Pipes proposed openstack/nova-specs master: Standardize CPU resource tracking https://review.openstack.org/555081 | |
| 15:55:28 | mriedem | dansmith: for the follow up https://review.openstack.org/#/c/572814/ | |
| 15:56:40 | dansmith | mriedem: don't we have to wait two years first? | |
| 15:56:47 | dansmith | like fine aged cheddar | |
| 15:57:04 | mriedem | artom linked a bz i don't have access to | |
| 16:00:00 | dansmith | mriedem: I looked and I don't think it has any useful context in it | |
| 16:00:54 | mriedem | ok | |
| 16:01:07 | mriedem | i'll proceed with todo cleanupification | |
| 16:02:20 | openstackgerrit | Merged openstack/nova master: Fix doc nit https://review.openstack.org/572682 | |
| 16:02:44 | openstackgerrit | Curt Moore proposed openstack/nova-specs master: Add spec for downloading images via RBD https://review.openstack.org/572805 | |
| 16:03:47 | mriedem | johnthetubaguy: jaypipes: bauzas: the vmware live migration spec needs another +2 https://review.openstack.org/#/c/299207/ | |
| 16:04:03 | bauzas | mriedem: looking | |
| 16:05:41 | openstackgerrit | Curt Moore proposed openstack/nova-specs master: Add spec for downloading images via RBD https://review.openstack.org/572805 | |
| 16:07:09 | openstackgerrit | Merged openstack/nova-specs master: Update multiple fixed-IPs support with services field https://review.openstack.org/572814 | |
| 16:08:04 | bauzas | mriedem: in pre_live_migration() on the dest compute, we don't manage exceptions raised by the virt driver, right? | |
| 16:08:21 | bauzas | mriedem: I think we only do this kind of pre-flight check in the conductor layer | |
| 16:10:13 | efried | stephenfin: You around? I need some undumbification about tunnel providers. | |
| 16:10:20 | stephenfin | efried: shoot | |
| 16:10:30 | efried | stephenfin: https://review.openstack.org/#/c/564440/3/nova/conf/neutron.py | |
| 16:10:38 | mriedem | bauzas: pre_live_migration is an rpc call from the source compute | |
| 16:10:48 | mriedem | so yes if that fails, the source sets the migration as an error and stops | |
| 16:11:17 | mriedem | bauzas: https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L6077 | |
| 16:11:32 | efried | stephenfin: What I thought I saw there is that you can only have your tunnel provider (singular) associated with one NUMA node. Is that right? | |
| 16:11:50 | stephenfin | efried: Nope | |
| 16:12:03 | stephenfin | I'm using the term tunnel provider but what I really mean is tunnel endpoint | |
| 16:12:17 | bauzas | mriedem: ok, so that will work, because I wasn't seeing any exception coming from the virt layer handled by the dest host | |
| 16:12:26 | stephenfin | namely, the logical NIC where all tunneled traffic ingresses and egresses from the compute node | |
| 16:12:42 | bauzas | mriedem: the spec is a bit missing of details, but okay | |
| 16:12:44 | stephenfin | logical because it can actually be two or more physical NICs bonded together | |
| 16:13:02 | efried | stephenfin: Ah. And those multiple NICs may live on multiple NUMA nodes. | |
| 16:13:07 | openstackgerrit | Murali Annamneni proposed openstack/nova master: Enables MySQL Cluster Support for Nova https://review.openstack.org/446643 | |
| 16:13:13 | stephenfin | you're only allowed one tunnel endpoint in neutron. Every tunneled network must use that endpoint | |
| 16:13:18 | stephenfin | efried: Correct | |
| 16:13:34 | stephenfin | Ditto for physnet-type networks (e.g. layer 2) | |
| 16:14:08 | stephenfin | You could have bonded NICs spread across the various host NUMA nodes | |
| 16:14:25 | efried | k, I get that. | |
| 16:14:48 | efried | Thanks stephenfin. | |
| 16:15:10 | stephenfin | no problemo | |
| 16:21:57 | jaypipes | mriedem: +W | |
| 16:22:35 | bauzas | jaypipes: snap, I litterally +Wd 10 secs after | |
| 16:22:39 | bauzas | but I have comments on it | |
| 16:26:11 | mriedem | bauzas: i replied to your comments about pre live migratoin | |
| 16:26:16 | mriedem | i don't really understand the concern | |
| 16:27:19 | bauzas | mriedem: it's a pure implementation thing | |
| 16:27:41 | bauzas | mriedem: it's just the fact we pass straight to the compute the migrate_data object that we build by the conductor | |
| 16:27:43 | bauzas | or the api | |
| 16:27:47 | stephenfin | efried: You're not on openstack-oslo but https://review.openstack.org/#/c/564822/ | |
| 16:27:52 | mriedem | "you somehow need to hook into this and call the virt driver so that you can raise the exception." | |
| 16:27:56 | mriedem | bauzas: ^ is all over rpc call | |
| 16:28:05 | mriedem | so if the virt driver raises an exception, it goes back to the source and fails | |
| 16:28:10 | bauzas | mriedem: but the source host never touches it *before* the dest compute amends it | |
| 16:28:26 | stephenfin | efried: I spotted the same issue on oslo.logging (I think) earlier this week | |