Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-05
16:01:17 efried avolkov Yeah, I get it. If they're different, then it makes more sense for each PF to have its own label, and you do your HA/anti-affinity by constructing your flavors appropriately.
16:01:17 avolkov efried: with original properties you can ask distinct ports for one boot request and distict switches for another
16:02:22 efried Problem with that, though, is if you don't have exactly the same number and configuration of SR-IOV cards on all your hosts.
16:02:39 edmondsw will it be possible to request ports on separate identical cards (HA if a card fails)?
16:03:05 efried So it probably makes more sense to *call* them groups anyway, even if they usually/always map to PFs :)
16:03:15 sean-k-mooney efried: i would prefer if we did not model this in flavor or image properties though
16:03:20 openstackgerrit Matt Riedemann proposed openstack/nova stable/pike: Cleanup allocations on invalid dest node during live migration https://review.openstack.org/500908
16:03:20 openstackgerrit Matt Riedemann proposed openstack/nova stable/pike: Add functional recreate test for live migration pre-check fails https://review.openstack.org/500907
16:03:28 efried sean-k-mooney Which "this"?
16:03:36 sean-k-mooney efried: really we should try to model the bond requiremetn as an atribute of the neutron port
16:03:46 sean-k-mooney vf selection policy
16:04:25 avolkov sean-k-mooney: +1 not to use flavors )
16:04:38 efried Yeah, that makes sense, sorry.
16:04:54 openstackgerrit melanie witt proposed openstack/nova master: Request zero root disk for boot-from-volume instances https://review.openstack.org/428481
16:04:55 openstackgerrit melanie witt proposed openstack/nova master: Claim and report zero root disk for boot-from-volume instances https://review.openstack.org/428505
16:05:06 efried but wait
16:05:25 efried wouldn't we like to be able to do a spawn with SR-IOV VFs in one command rather than two?
16:05:33 sean-k-mooney efried: did you see the section i added to the ptg etherpad https://etherpad.openstack.org/p/nova-ptg-queens lines 107-125
16:05:40 sean-k-mooney efried: no
16:06:01 sean-k-mooney efried: and yes but not via flavor
16:06:12 efried Is a port bound to a host?
16:06:38 sean-k-mooney if we can do it with one command via nova-boot sure but i dont whant to create a flavor multiple time with just different number of interfaes
16:06:55 sean-k-mooney efried: it is after nova selects the host
16:07:10 sean-k-mooney efried: before that it is just a logical port in a db
16:07:32 sean-k-mooney efried: as part of portbinding nova compute updates the neutron port with the host id
16:08:05 efried So you want to create the port with an anti-affinity/HA spec which specifies the number of VFs and how they should be spread out...
16:08:25 sean-k-mooney efried: yep
16:08:49 openstackgerrit Matt Riedemann proposed openstack/nova master: Skip more racy rebuild failing tests with cells v1 https://review.openstack.org/499001
16:08:55 efried ...and then we schedule to the host and spawn attaches the right number of VFs with the right distribution.
16:10:02 efried There's a big hand-wavey part in the middle there, though, where the scheduler was able to figure out which compute host(s) would be able to honor that request. Is the scheduler (and/or, gods forbid, the placement API) supposed to introspect the port metadata to help with that decision??
16:10:14 sean-k-mooney efried: yep see my comments in https://review.openstack.org/#/c/463526/ and https://review.openstack.org/#/c/182242/ to this effect
16:11:32 sean-k-mooney no basically before the scheduler starts scheduling today the neutron v2 client api in nova retrives the port from neutorn
16:12:06 sean-k-mooney if that port is vnic_type direct/macvtap or virtio-forwarder it create a new pcieresutespec object
16:12:49 sean-k-mooney we need to extend that to also read the ha spec and add that to the picerequeste spec so that when tha tis passed to the sceduler/placement it can fufille the requirementes for ha
16:13:19 sean-k-mooney this is what we have imlemented for the feature based scheduing also
16:13:20 efried Yeah, okay, so that's how it works today; but I thought we were trying to move away from that kind of special-casing as we get into placement.
16:13:48 sean-k-mooney efried: yes so in placement i would like to be able to express affinity and anti affintiy
16:14:02 sean-k-mooney so we would ask for 2 vf with pf antiaffinity
16:14:27 sean-k-mooney placement would filer host based on that and then the scheduler would make the final desision
16:15:28 efried Has jaypipes weighed in yet on how affinity/anti-affinity might be made to work with placement?
16:16:09 dansmith efried: distance
16:16:13 sean-k-mooney proably its not the first time i have mentioned this to him but not aware of his current stance
16:16:22 dansmith efried: as mentioned earlier this morning, but also quite a bit in boston during that session
16:16:53 edmondsw there are anti-affinity needs at multiple layers... 1) device, 2) PF, 3) switch...
16:17:24 sean-k-mooney edmondsw: yes i was hoping we could model the switch as a trait on the pf if that made sense?
16:17:41 efried Right, so basically placement would need a generic, multi-layer-capable distance/affinity mechanism, and then consumers could model as they see fit within that framework.
16:18:32 sean-k-mooney efried: yes ideally. its just up to the consumers to model the dependcies with traits and netested providers correctly
16:18:41 sean-k-mooney that is easier said then done however
16:19:00 sean-k-mooney ideally i would like to see a request for a bonded port that looked someting like this
16:19:01 efried sean-k-mooney Definitely makes sense for the switch to be a trait on the PF. Or for the switch to be a RP with its PFs nested underneath it. Either way would work. But if these affinity gizmos are separate from traits, then it probably doesn't matter as much how the RPs are nested.
16:19:02 sean-k-mooney neutron port-create --binding:vnic_type=direct --binding:profile={bond=true, bond_mode=active-backup,bound_count=2,bond_antiafinity=pf} --name bond1 private
16:19:36 efried Where you could conceivably also say bond_antiaffinity=pf,switch ?
16:20:01 sean-k-mooney efried: yes
16:20:14 jaypipes efried: distance would be stored in an aggregate_distances table, which would store the distance between the providers in one aggregate and providers in another.
16:20:31 jaypipes efried: distance would just be a number. higher the number, greater the relative distance.
16:21:18 edmondsw or bond_antiaffinity=card,switch if you want the PFs to come from different physical CNAs?
16:21:40 sean-k-mooney jaypipes: im not sure that distance would be a good match to antiafinity/affinity if i need to manually create agggreates for the vfs but ingeneral it is usefull
16:22:34 sean-k-mooney edmondsw: you could but you see the general pattern i would love to express this requirement on the neutorn port in a general way rather then statically defiened in a flavor
16:23:01 edmondsw sean-k-mooney yeah, and I think I agree with that if we can make it work
16:23:03 efried Now, let me make sure I understand something: placement is never going to be in the business of assigning individual VFs around. It's just gonna decrement the count and say "you got one from this RP".
16:23:04 sean-k-mooney bassically i would like to keep flavor extraspecs for compute requirement
16:23:59 efried Then it's up to the virt driver (or maybe the mech driver in this case?) to decide which specific VF to use -- or even to create the VF on the fly if that's something it can do.
16:24:03 sean-k-mooney efried: maybe... i think jay would agree i would not mind extending it to be able to do indvidual assignment
16:24:24 jaypipes sean-k-mooney: it's a possibility.
16:24:48 efried Point is, placement shouldn't be aware of an individual VF any more than it's aware of an individual memory megabyte.
16:25:13 jaypipes sean-k-mooney: in the same way that we agreed not to have aggregates have traits (instead, we "push down" all traits to the resource provider)
16:25:26 jaypipes efried: not necessarily.
16:25:50 dansmith jaypipes: eh?
16:25:55 efried jaypipes: eh?
16:26:00 jaypipes efried: if you need to differentiate between two VFs on a host because those two VFs expose different capabilities, then you will need to create each different VF as a resource provider.
16:26:01 sean-k-mooney efried: well i would like to have the mem_page resouce provider track indiviual pages too but i have more importing things to adress first
16:26:12 dansmith ah, sure
16:26:16 efried Okay, yeah.
16:26:21 dansmith not sure when/how that would happen
16:26:26 dansmith but if it did, then I guess
16:26:37 dansmith although then we're going to have a shitton of single-resource providers
16:26:40 jaypipes dansmith: ask sean-k-mooney. Intel excels at creating uses for complexity.
16:27:16 dansmith I'd hope that you could separate that into 32 VFs with tls-offload and 32 without, for a 64-vf nic
16:27:18 dansmith but..
16:27:18 efried okay, as long as the general case is to have the inventory of VFs just be a number.
16:27:29 dansmith efried: it's still that,
16:27:35 dansmith efried: you'd just have inventory=1 for these
16:27:39 efried yeah yeah.
16:27:42 efried I get it.
16:27:55 efried But the tls-offload thing...
16:28:00 sean-k-mooney well i can certenly do a lot with singel-resouce proverders
16:28:03 efried I would really hope there would be some other way to specify that.
16:28:15 dansmith efried: I wouldn't
16:28:23 jaypipes sean-k-mooney: would single resource providers provide spell checking? :)
16:28:43 sean-k-mooney haha perhaps...
16:28:44 dansmith efried: tls-offload being a trait for regular nics, and vfs, so if some have it and some don't...
16:28:48 jaypipes sean-k-mooney: :P
16:29:16 efried But if I create my VFs on the fly, and could assign that trait to any one of 'em on the fly (say, based on a prop in the binding profile)...
16:29:35 jaypipes dansmith, efried: technically you wouldn't necessarily need to have single resource providers for each VF. just one resource provider per set of VFs with similar capabilities.
16:29:37 dansmith efried: no, that's not the same
16:29:44 dansmith jaypipes: right exactly
16:30:00 dansmith efried: I meant if some VFs would be unable to provide tls offload, then they go in a separate provider without that trait
16:30:28 sean-k-mooney so our current generation of nics cant do this but in future nics we will be able to load firmware that gives different feature per vf
16:30:34 dansmith efried: if they all can and it's a by-request thing, then that means they're all in one provider with that trait and the virt driver decides to configure it as such based on the request
16:30:41 sean-k-mooney with fortvile XL710 its card wide

Earlier   Later