| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-25 | |||
| 13:58:03 | openstackgerrit | Ed Leafe proposed openstack/nova-specs master: Return Alternate Hosts https://review.openstack.org/504275 | |
| 13:58:09 | jaypipes | sahid: but anyway, like I said, I'm fine with you adding to the existing code. | |
| 13:58:21 | jaypipes | sahid: I just will be focusing on the n-r-p stuff, that's all | |
| 13:58:40 | bauzas | it's open-source development, anyone can just propose | |
| 13:58:58 | sahid | jaypipes: yes i understand that which it make sense and i will be happy to help and review some part if i can | |
| 13:59:03 | bauzas | I just felt we discussed on what could be achievable for Queens in parallel of the nested RPs implementation | |
| 14:00:02 | bauzas | the main focus for Queens was to provide the libvirt feature to returning RPs that were augmented by the VGPU resources | |
| 14:00:38 | bauzas | jianghuaw: can you clarify how Xen would see those GPU types ? | |
| 14:00:54 | edleafe | Scheduler subteam meeting running now in #openstack-meeting-alt | |
| 14:00:55 | sahid | bauzas: so you want to push on top of something under heavy developpement a feature which we are going to be used by large industries ? i think we should be reasonable | |
| 14:01:28 | bauzas | sahid: I'm just saying it's orthogonal :) | |
| 14:01:33 | sahid | RP is doind a lof of things, we need a deprecating phase, where some users can migrate and so we can fix the issues | |
| 14:01:35 | jianghuaw | bauzas, ok. XenServer will detect the pGPU types and make the same type of pGPUs into a single group. | |
| 14:01:46 | bauzas | vGPU tracking is done by providing a new set of resource classes | |
| 14:01:53 | bauzas | and traits | |
| 14:02:11 | bauzas | while nested resource providers is focusing on providing a tree of resource providers with dependencies | |
| 14:02:19 | bauzas | those don't overlap | |
| 14:02:24 | jianghuaw | And expose the vGPU types supported by the pGPU group. | |
| 14:02:47 | jaypipes | sahid: for the record, we *have* been doing that deprecation/migration period with resource providers. for example, we had a deprecation/migration period for tracking Ironic nodes as atomic resources. | |
| 14:02:49 | bauzas | snap, scheduler meeting | |
| 14:03:03 | jianghuaw | when requesting a vGPU, xenserver will schedule a PGPU and create a vGPU on it. | |
| 14:03:14 | jaypipes | sahid: so it's not true that we're just imposing resource providers modeling without providing a migration path for existing resource classes. | |
| 14:03:41 | jianghuaw | bauzas, so we can't define whitelist per pGPU by per pGPU group. | |
| 14:03:51 | bauzas | sahid: jaypipes: probably a good call for discussing that in the scheduler meeting | |
| 14:03:54 | sahid | jaypipes: oh yes right, but, what about NUMA, CPU Pinning, Huge Pages and all the virt specific features? | |
| 14:04:08 | sahid | don't we have to migr incrementally? | |
| 14:04:18 | jaypipes | sahid: yes, absolutely. | |
| 14:04:32 | jaypipes | sahid: none of those resources are planned to migrate in Queens, BTW. | |
| 14:04:41 | sahid | jaypipes: yes and what is the ETA for sometging production-ready? | |
| 14:04:48 | sahid | ok cool that is my point | |
| 14:04:50 | bauzas | well, with the fact that CPU pinning isn't targeted to be pushed to Placement, right? | |
| 14:04:50 | owalsh | bauzas: hey | |
| 14:04:53 | jianghuaw | bauzas, that's why we make the pGPU group as the vGPU's resource provider. | |
| 14:05:16 | bauzas | jianghuaw: so, say I have two same cards, I'd only see one pGPU group right? | |
| 14:05:27 | jianghuaw | true. | |
| 14:05:28 | sahid | jaypipes: it will take at least 2 or 3 releases and during that time we can't make any development? | |
| 14:05:39 | bauzas | jianghuaw: okay, then that's a bit different from libvirt | |
| 14:06:01 | sahid | i just suggest we work in parallel | |
| 14:06:29 | jianghuaw | bauzas, could you give an example for 'that shows up different device IDs for two same cards.' | |
| 14:06:32 | sahid | since the /pci is kind of production-ready and provide eveyrything we need | |
| 14:06:32 | jaypipes | sahid: depends on what you mean by "production-ready". (I personally don't think the existing pci manager is well-written or maintainable, but I guess there are different defninitions of "production-ready") | |
| 14:06:50 | jaypipes | sahid: again, I'm not disagreeing with you... | |
| 14:06:56 | sahid | jaypipes: yes yes :) | |
| 14:07:00 | jaypipes | not sure why you're acting like I am :) | |
| 14:07:26 | sahid | oops sorry really | |
| 14:07:45 | sahid | it's just that you are the only who are paying attention at me :) | |
| 14:08:10 | dansmith | jaypipes: we can provide pretty simple vgpu support via resource classes today right? without building more into the pci infrastructure we have | |
| 14:08:31 | jaypipes | sahid: we're paying attention but also in scheduler IRC meeting :) | |
| 14:08:42 | dansmith | jaypipes: if we just let virt drivers see that a vgpu was requested (i.e. see the flavor) and just configure the guest with an available one | |
| 14:08:43 | jianghuaw | bauzas, will the "GRID M60-0B" have different type id <type id='nvidia-11'> for two same pGPUs? | |
| 14:09:20 | bauzas | jianghuaw: I haven't tested yet, but I bet it | |
| 14:09:27 | jaypipes | dansmith: very simple resources, yes. we could add a (custom) resource class for the VGPU resources and have a flavor consume some amount of those | |
| 14:09:43 | dansmith | jaypipes: right, that'd be my preference for the first go-round | |
| 14:10:02 | jianghuaw | bauzas, I think if the type id is different, it will be a good solution to use 'type-id' for libvirt. | |
| 14:10:20 | bauzas | dansmith: jaypipes: I was only seeing the virt drivers adding the new VGPU resource class to the existing RPs as a first round, that's it :) | |
| 14:10:26 | dansmith | jaypipes: let people segregate different types of gpus into aggregates, and just assume equal portions of vgpu per instance without any smarts, which I think would be a perfectly reasonable first stab, which we can iterate on when we have traits and things | |
| 14:10:32 | dansmith | bauzas: ++ | |
| 14:11:06 | jaypipes | dansmith: agreed. | |
| 14:11:21 | bauzas | jianghuaw: yeah, so you agree with the fact that the listOpt can have different semantics based on the driver? | |
| 14:11:50 | jianghuaw | bauzas, yes. libvirt uses the type-id but XenServer continue to use the type name. As you suggested, virt driver should able to mapping them into real resources. | |
| 14:12:05 | bauzas | that would be strings for each of the drivers, but in the libvirt case, it would be device IDs while it would be pGPU resource type names for xen | |
| 14:12:08 | jianghuaw | but we need confirm if the type-id will be different. | |
| 14:12:19 | bauzas | jianghuaw: what I need is a testbed :p | |
| 14:13:02 | jianghuaw | bauzas, thanks. | |
| 14:13:28 | bauzas | jianghuaw: fancy updating the spec so I can do a pass once again ? | |
| 14:14:10 | jianghuaw | bauzas, sure. I will update it soon. | |
| 14:14:19 | bauzas | cool | |
| 14:14:28 | sahid | dansmith: basically you can't do more than that right now and it's exactly what i'm pointing where we could have better support | |
| 14:14:36 | jianghuaw | bauzas, so I think the option name as enabled_vgpu_types is fine. right? | |
| 14:14:58 | jianghuaw | just confirm if I can keep this option name:-) | |
| 14:15:25 | celebrate | Hi guys! Do you know where the "availability_zone" attribute of instance is stored? | |
| 14:15:28 | dansmith | sahid: building more on the current pci stuff isn't "better" in any way, IMHO :) | |
| 14:15:57 | dansmith | sahid: and I just think that providing a basic building block initially and moving on from there avoids us building more baggage that we have to migrate, adapt, and remove later | |
| 14:16:00 | mriedem | celebrate: instance.availability_zone, or as metadata on a host aggregate the instance is in | |
| 14:16:25 | bauzas | celebrate: the scheduler queries hosts based on request_spec.availability_zone | |
| 14:16:41 | bauzas | instance.az is only set at boot time | |
| 14:16:50 | celebrate | yeah, in database i see AZ=nova, but scheduler thinks its "None" | |
| 14:17:12 | sahid | dansmith: i'm not agree with you the /pci is good and do what we need | |
| 14:17:27 | dansmith | sahid: okay :) | |
| 14:17:35 | sahid | dansmith: in all cases we will have to migrate the /pci | |
| 14:17:44 | celebrate | Do you know how to make scheduler understand that AZ is "nova"? | |
| 14:17:58 | mriedem | sahid: we don't have to migrate vgpus if we don't implement it in /pci | |
| 14:17:59 | sahid | the support of mdev is not so heavy | |
| 14:18:12 | sahid | mriedem: mdev is just pci devices | |
| 14:18:14 | bauzas | celebrate: well, nova AZ is actually equal to None | |
| 14:18:29 | sahid | so when you will migrate sriov it will be the same work | |
| 14:18:50 | bauzas | celebrate: that's a convention, or rather a config option default value | |
| 14:19:08 | mriedem | celebrate: https://docs.openstack.org/nova/latest/user/aggregates.html | |
| 14:19:11 | bauzas | namely [DEFAULT]/default_availability_zone | |
| 14:19:37 | celebrate | I've already set [DEFAULT]/default_availability_zone to "nova" | |
| 14:19:56 | celebrate | however scheduler thinks its None and refuses to respect AZ from aggregates | |
| 14:20:15 | bauzas | celebrate: I will need to disappear in a couple of minutes, but here are a few things to know | |
| 14:20:41 | bauzas | celebrate: #1 if you didn't ask a specific AZ at boot time, that means the scheduler won't care about placing your instance to a specific AZ | |
| 14:21:19 | bauzas | celebrate: if you did specified a AZ at boot, AZFilter in the scheduler will check the AZ for each of the hosts | |
| 14:21:28 | celebrate | can I somehow make scheduler change its mind? | |
| 14:21:38 | bauzas | celebrate: not sure I get why :) | |
| 14:22:07 | celebrate | because I have different AZ and scheduler sends VM there during live-migration | |
| 14:22:33 | celebrate | where this information is stored in DB? | |
| 14:23:03 | bauzas | celebrate: could you please describe your problem ? | |
| 14:23:36 | celebrate | I have 3 different AZ and during live-migration VMs travel from one AZ to another | |