| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-05 | |||
| 15:51:43 | mriedem | 3pm wednesday would be *after* lunch :) | |
| 15:52:52 | edmondsw | group 1 = one port to sw1 and one to sw2, group 2 = one port to sw1 and one to sw2, then ask for 2 ports in each group | |
| 15:53:49 | efried | edmondsw But placement doesn't know enough to not give you both VFs from group 1 on the same switch. | |
| 15:54:18 | edmondsw | efried there aren't 2 ports in group 1 with the same switch | |
| 15:54:35 | efried | Yes, there are multiple VFs on each pport. | |
| 15:54:44 | edmondsw | oh, VFs... | |
| 15:55:06 | edmondsw | I gotcha now | |
| 15:55:07 | efried | And yes, you could conceivably do this same thing with four groups - but enumerating switches might make more sense to the user; and you also want the model to extend to >2 switches, >2 ports per switch. | |
| 15:55:29 | dtantsur | mriedem: okay, sounds good | |
| 15:56:16 | efried | although... avolkov that might actually make more sense. If you always know you want four VFs, you could just tag your PFs in groups so they'll always be spread out. | |
| 15:56:20 | edmondsw | efried how about group 1 = 2 ports to sw1 and group 2 = 2 ports to sw2? | |
| 15:56:55 | efried | edmondsw Then again you'll ask for two VFs from group 1 and they might wind up coming from the same pport | |
| 15:57:17 | edmondsw | yep, k... better for PFs but doesn't help with VFs | |
| 15:57:30 | efried | eh? | |
| 15:57:41 | edmondsw | nm... it doesn't work, so it doesn't work :) | |
| 15:58:16 | edmondsw | I'm not following why you wouldn't tag them with the port, then | |
| 15:58:29 | efried | avolkov In your email example, it's no different than having labeled P1,P2,P3,P4 - but extending to more than four pports (or reducing the problem set to fewer than 4 desired VFs) it makes more sense to think of groups - where the total number of groups is the number of VFs you're going to want from a single allocation request. | |
| 15:58:34 | efried | edmondsw ^^ | |
| 16:00:12 | avolkov | efried: groups are okay if you have the same requirements for each boot request | |
| 16:01:17 | avolkov | efried: with original properties you can ask distinct ports for one boot request and distict switches for another | |
| 16:01:17 | efried | avolkov Yeah, I get it. If they're different, then it makes more sense for each PF to have its own label, and you do your HA/anti-affinity by constructing your flavors appropriately. | |
| 16:02:22 | efried | Problem with that, though, is if you don't have exactly the same number and configuration of SR-IOV cards on all your hosts. | |
| 16:02:39 | edmondsw | will it be possible to request ports on separate identical cards (HA if a card fails)? | |
| 16:03:05 | efried | So it probably makes more sense to *call* them groups anyway, even if they usually/always map to PFs :) | |
| 16:03:15 | sean-k-mooney | efried: i would prefer if we did not model this in flavor or image properties though | |
| 16:03:20 | openstackgerrit | Matt Riedemann proposed openstack/nova stable/pike: Add functional recreate test for live migration pre-check fails https://review.openstack.org/500907 | |
| 16:03:20 | openstackgerrit | Matt Riedemann proposed openstack/nova stable/pike: Cleanup allocations on invalid dest node during live migration https://review.openstack.org/500908 | |
| 16:03:28 | efried | sean-k-mooney Which "this"? | |
| 16:03:36 | sean-k-mooney | efried: really we should try to model the bond requiremetn as an atribute of the neutron port | |
| 16:03:46 | sean-k-mooney | vf selection policy | |
| 16:04:25 | avolkov | sean-k-mooney: +1 not to use flavors ) | |
| 16:04:38 | efried | Yeah, that makes sense, sorry. | |
| 16:04:54 | openstackgerrit | melanie witt proposed openstack/nova master: Request zero root disk for boot-from-volume instances https://review.openstack.org/428481 | |
| 16:04:55 | openstackgerrit | melanie witt proposed openstack/nova master: Claim and report zero root disk for boot-from-volume instances https://review.openstack.org/428505 | |
| 16:05:06 | efried | but wait | |
| 16:05:25 | efried | wouldn't we like to be able to do a spawn with SR-IOV VFs in one command rather than two? | |
| 16:05:33 | sean-k-mooney | efried: did you see the section i added to the ptg etherpad https://etherpad.openstack.org/p/nova-ptg-queens lines 107-125 | |
| 16:05:40 | sean-k-mooney | efried: no | |
| 16:06:01 | sean-k-mooney | efried: and yes but not via flavor | |
| 16:06:12 | efried | Is a port bound to a host? | |
| 16:06:38 | sean-k-mooney | if we can do it with one command via nova-boot sure but i dont whant to create a flavor multiple time with just different number of interfaes | |
| 16:06:55 | sean-k-mooney | efried: it is after nova selects the host | |
| 16:07:10 | sean-k-mooney | efried: before that it is just a logical port in a db | |
| 16:07:32 | sean-k-mooney | efried: as part of portbinding nova compute updates the neutron port with the host id | |
| 16:08:05 | efried | So you want to create the port with an anti-affinity/HA spec which specifies the number of VFs and how they should be spread out... | |
| 16:08:25 | sean-k-mooney | efried: yep | |
| 16:08:49 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Skip more racy rebuild failing tests with cells v1 https://review.openstack.org/499001 | |
| 16:08:55 | efried | ...and then we schedule to the host and spawn attaches the right number of VFs with the right distribution. | |
| 16:10:02 | efried | There's a big hand-wavey part in the middle there, though, where the scheduler was able to figure out which compute host(s) would be able to honor that request. Is the scheduler (and/or, gods forbid, the placement API) supposed to introspect the port metadata to help with that decision?? | |
| 16:10:14 | sean-k-mooney | efried: yep see my comments in https://review.openstack.org/#/c/463526/ and https://review.openstack.org/#/c/182242/ to this effect | |
| 16:11:32 | sean-k-mooney | no basically before the scheduler starts scheduling today the neutron v2 client api in nova retrives the port from neutorn | |
| 16:12:06 | sean-k-mooney | if that port is vnic_type direct/macvtap or virtio-forwarder it create a new pcieresutespec object | |
| 16:12:49 | sean-k-mooney | we need to extend that to also read the ha spec and add that to the picerequeste spec so that when tha tis passed to the sceduler/placement it can fufille the requirementes for ha | |
| 16:13:19 | sean-k-mooney | this is what we have imlemented for the feature based scheduing also | |
| 16:13:20 | efried | Yeah, okay, so that's how it works today; but I thought we were trying to move away from that kind of special-casing as we get into placement. | |
| 16:13:48 | sean-k-mooney | efried: yes so in placement i would like to be able to express affinity and anti affintiy | |
| 16:14:02 | sean-k-mooney | so we would ask for 2 vf with pf antiaffinity | |
| 16:14:27 | sean-k-mooney | placement would filer host based on that and then the scheduler would make the final desision | |
| 16:15:28 | efried | Has jaypipes weighed in yet on how affinity/anti-affinity might be made to work with placement? | |
| 16:16:09 | dansmith | efried: distance | |
| 16:16:13 | sean-k-mooney | proably its not the first time i have mentioned this to him but not aware of his current stance | |
| 16:16:22 | dansmith | efried: as mentioned earlier this morning, but also quite a bit in boston during that session | |
| 16:16:53 | edmondsw | there are anti-affinity needs at multiple layers... 1) device, 2) PF, 3) switch... | |
| 16:17:24 | sean-k-mooney | edmondsw: yes i was hoping we could model the switch as a trait on the pf if that made sense? | |
| 16:17:41 | efried | Right, so basically placement would need a generic, multi-layer-capable distance/affinity mechanism, and then consumers could model as they see fit within that framework. | |
| 16:18:32 | sean-k-mooney | efried: yes ideally. its just up to the consumers to model the dependcies with traits and netested providers correctly | |
| 16:18:41 | sean-k-mooney | that is easier said then done however | |
| 16:19:00 | sean-k-mooney | ideally i would like to see a request for a bonded port that looked someting like this | |
| 16:19:01 | efried | sean-k-mooney Definitely makes sense for the switch to be a trait on the PF. Or for the switch to be a RP with its PFs nested underneath it. Either way would work. But if these affinity gizmos are separate from traits, then it probably doesn't matter as much how the RPs are nested. | |
| 16:19:02 | sean-k-mooney | neutron port-create --binding:vnic_type=direct --binding:profile={bond=true, bond_mode=active-backup,bound_count=2,bond_antiafinity=pf} --name bond1 private | |
| 16:19:36 | efried | Where you could conceivably also say bond_antiaffinity=pf,switch ? | |
| 16:20:01 | sean-k-mooney | efried: yes | |
| 16:20:14 | jaypipes | efried: distance would be stored in an aggregate_distances table, which would store the distance between the providers in one aggregate and providers in another. | |
| 16:20:31 | jaypipes | efried: distance would just be a number. higher the number, greater the relative distance. | |
| 16:21:18 | edmondsw | or bond_antiaffinity=card,switch if you want the PFs to come from different physical CNAs? | |
| 16:21:40 | sean-k-mooney | jaypipes: im not sure that distance would be a good match to antiafinity/affinity if i need to manually create agggreates for the vfs but ingeneral it is usefull | |
| 16:22:34 | sean-k-mooney | edmondsw: you could but you see the general pattern i would love to express this requirement on the neutorn port in a general way rather then statically defiened in a flavor | |
| 16:23:01 | edmondsw | sean-k-mooney yeah, and I think I agree with that if we can make it work | |
| 16:23:03 | efried | Now, let me make sure I understand something: placement is never going to be in the business of assigning individual VFs around. It's just gonna decrement the count and say "you got one from this RP". | |
| 16:23:04 | sean-k-mooney | bassically i would like to keep flavor extraspecs for compute requirement | |
| 16:23:59 | efried | Then it's up to the virt driver (or maybe the mech driver in this case?) to decide which specific VF to use -- or even to create the VF on the fly if that's something it can do. | |
| 16:24:03 | sean-k-mooney | efried: maybe... i think jay would agree i would not mind extending it to be able to do indvidual assignment | |
| 16:24:24 | jaypipes | sean-k-mooney: it's a possibility. | |
| 16:24:48 | efried | Point is, placement shouldn't be aware of an individual VF any more than it's aware of an individual memory megabyte. | |
| 16:25:13 | jaypipes | sean-k-mooney: in the same way that we agreed not to have aggregates have traits (instead, we "push down" all traits to the resource provider) | |
| 16:25:26 | jaypipes | efried: not necessarily. | |
| 16:25:50 | dansmith | jaypipes: eh? | |
| 16:25:55 | efried | jaypipes: eh? | |
| 16:26:00 | jaypipes | efried: if you need to differentiate between two VFs on a host because those two VFs expose different capabilities, then you will need to create each different VF as a resource provider. | |
| 16:26:01 | sean-k-mooney | efried: well i would like to have the mem_page resouce provider track indiviual pages too but i have more importing things to adress first | |
| 16:26:12 | dansmith | ah, sure | |
| 16:26:16 | efried | Okay, yeah. | |
| 16:26:21 | dansmith | not sure when/how that would happen | |
| 16:26:26 | dansmith | but if it did, then I guess | |
| 16:26:37 | dansmith | although then we're going to have a shitton of single-resource providers | |
| 16:26:40 | jaypipes | dansmith: ask sean-k-mooney. Intel excels at creating uses for complexity. | |
| 16:27:16 | dansmith | I'd hope that you could separate that into 32 VFs with tls-offload and 32 without, for a 64-vf nic | |
| 16:27:18 | efried | okay, as long as the general case is to have the inventory of VFs just be a number. | |
| 16:27:18 | dansmith | but.. | |
| 16:27:29 | dansmith | efried: it's still that, | |