| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-11-28 | |||
| 14:09:42 | jianghuaw_ | the amount of display heads is determined by the vgpu type. | |
| 14:10:04 | jianghuaw_ | shouldn't. | |
| 14:10:25 | efried | jianghuaw_ Then yeah, I say leave display heads out of the picture entirely. Don't register them as resources, and don't include them in the flavor. | |
| 14:11:12 | sdague | efried: +A | |
| 14:11:18 | efried | sdague Thanks! | |
| 14:11:51 | sdague | and +2 on the config matrix update | |
| 14:12:13 | efried | sdague Cool. | |
| 14:12:47 | jianghuaw_ | efried, but I guess some customers may want to request vGPU with needed heads. | |
| 14:14:28 | jianghuaw_ | jaypipes, efried: I'm feeling the display heads somehow looks like a trait. | |
| 14:15:12 | jaypipes | jianghuaw_: it's not :) it's a consumable resource | |
| 14:15:39 | efried | Unless it's like DUAL_DISPLAY_HEAD_CAPABLE :) | |
| 14:16:06 | jianghuaw_ | jaypipes, but it does have problem if we allow having display head as optional. | |
| 14:16:31 | edleafe | jianghuaw_: can you request a vGPU *without* consuming a display head? | |
| 14:16:40 | jaypipes | jianghuaw_: why is that? wouldn't you just have a flavor that doesn't have display heads? | |
| 14:17:46 | jianghuaw_ | efried, some vGPU types don't support display heads. | |
| 14:19:27 | jianghuaw_ | As some vGPUs don't support display heads. so the display heads' inv may be empty. So we should allow the flavor without specifying display heads. | |
| 14:20:50 | efried | jianghuaw_ In those cases, it makes sense. Where I think we're having issues is this idea that, by requesting a certain type of vGPU, we might "automatically" consume some number of display heads. We can't do that if we're inventorying the display heads. | |
| 14:20:51 | jianghuaw_ | But it's still possible to allocate a vGPU which support display heads it only request some kind of vGPU. | |
| 14:21:44 | efried | jianghuaw_ Let me ask it this way: In these cases where certain vGPU types support consuming display heads, is it possible to request a vGPU *without* consuming display heads? | |
| 14:21:46 | jianghuaw_ | efried, yes. that's exactly the problem I met. | |
| 14:22:19 | edleafe | efried: jinx-ish | |
| 14:22:51 | efried | jianghuaw_ If they need to be inventoried as separate resources, then they need to be requested as separate resources. | |
| 14:23:21 | jaypipes | efried: +1 | |
| 14:24:26 | efried | jianghuaw_ And it sounds like we may have cases where a particular flavor setup is simply invalid: If there's a vGPU type that *must* consume display heads, your flavor *must* request the appropriate number of vGPUs *and* display heads. Otherwise it's bogus. | |
| 14:25:03 | bauzas | jianghuaw_: sorry was afk | |
| 14:25:22 | jianghuaw_ | bauzas, hi. no worries. | |
| 14:25:26 | bauzas | oh, you're discussing about display heads ? | |
| 14:25:40 | jianghuaw_ | bauzas, yes. | |
| 14:25:47 | bauzas | ok | |
| 14:25:54 | mriedem | esberglu: efried: why are evacuate and rebuild marked as missing in here? https://review.openstack.org/#/c/523140/ | |
| 14:26:08 | mriedem | nova-compute has a default rebuild impl if the virt driver doesn't implement it's own | |
| 14:26:32 | bauzas | so, the point is that given some VGPU types don't provide display heads, then the inventory for them would be none | |
| 14:27:11 | bauzas | so, when asking for display heads, the placement API wouldn't provide the computes having them | |
| 14:27:14 | jianghuaw_ | yes. I think we have cases that the vGPUs don't support display heads co-exist with vGPU not supporting it. | |
| 14:27:45 | bauzas | then, the scheduler would allocate a claim only for the computes supporting them | |
| 14:27:56 | efried | mriedem We have logic to make sure devices are re-attached in the proper order on rebuild. That logic isn't in the in-tree driver yet. So a rebuild *might* work with the default impl, but it might not. | |
| 14:28:02 | jianghuaw_ | that's why I'm thinking we should allow flavors which don't request display heads. | |
| 14:28:03 | bauzas | jianghuaw_: see the problem ? | |
| 14:28:10 | bauzas | right | |
| 14:28:27 | mriedem | efried: that's why it was marked 'unknown' | |
| 14:29:00 | mriedem | efried: i don't think libvirt ensures devices are re-attached in the same order on rebuild either | |
| 14:29:21 | mriedem | that's why mdbooth had the bdm metadata stuff for awhile | |
| 14:29:25 | mriedem | *i think* | |
| 14:30:08 | mdbooth | efried: They are definitely *not* reattached in the same order in all circumstances | |
| 14:30:24 | jianghuaw_ | bauzas, It's possible to allocate vGPUs which actually supports display heads for those flavors not requesting display heads. | |
| 14:30:27 | efried | mdbooth Yeah, so in PowerVM, we guarantee that. | |
| 14:30:55 | jianghuaw_ | as the resource provider have inventory for VGPU which can meet the request. | |
| 14:31:07 | mdbooth | efried: It would be nice. | |
| 14:31:37 | mdbooth | efried: Is your scheme reproducible? | |
| 14:31:39 | efried | mriedem Yeah, I confirmed by looking at our OOT driver - there's a bunch of special stuff we do in spawn for the rebuild case that we haven't ported in yet. So we want to declare that we *don't* support rebuild in the in-tree driver yet. | |
| 14:31:59 | bauzas | jianghuaw_: the problem I see is that if you provide a flavor asking for both VGPUs *and* heads, then the scheduler will ask placement for *both* | |
| 14:32:03 | mdbooth | efried: i.e. Could the libvirt driver use it? | |
| 14:32:26 | efried | mdbooth Bunch of caveats, but I think the concept might be portable. | |
| 14:32:42 | bauzas | jianghuaw_: AFAIR, we had a spec for modifying the placement call to ask for some resource classes not mandatory | |
| 14:32:51 | bauzas | efried: jaypipes: is it, right? | |
| 14:33:03 | efried | bauzas Not for Q | |
| 14:33:04 | mdbooth | The issue the libvirt driver has is the interaction between dynamic and 'static' disks | |
| 14:33:12 | bauzas | efried: yup, I know | |
| 14:33:25 | bauzas | efried: I was more commenting about the fact we discussed that | |
| 14:33:27 | efried | bauzas The one I remember was for "preferred traits", not for optional resource classes. | |
| 14:33:33 | bauzas | yeah that | |
| 14:33:37 | mdbooth | So if you have an instance with a root disk and a config disk, then you attach a volume, you'll have root disk, config disk, volume | |
| 14:33:49 | mdbooth | But in rebuild we always put the config disk last | |
| 14:33:58 | mdbooth | So on rebuild you'll have root disk, volume, config disk | |
| 14:33:59 | jianghuaw_ | bauzas, I think this case makes sense. What I'm think is that a flavor is asking only VGPU, but it placement may return RPs which contains display heads also. | |
| 14:34:10 | mdbooth | And in fact we have no way of knowing what the previous order was | |
| 14:34:13 | jianghuaw_ | it will result into the display heads leaking. | |
| 14:34:16 | bauzas | jianghuaw_: right | |
| 14:34:37 | bauzas | what do you mean by leaking ? | |
| 14:34:43 | efried | mdbooth For us, the dev order is determined by which virtual "slot" the disk is attached via. So we maintain a cache of which devices are associated with which slots, and then we make sure to restore that on rebuild. | |
| 14:34:50 | bauzas | jianghuaw_: that's depending on the virt driver, right ? | |
| 14:35:00 | mdbooth | efried: Right. Where do you store the cache, though? | |
| 14:35:15 | efried | mdbooth In our case, in order to support that, the user has to set up swift :$ | |
| 14:35:27 | mdbooth | Ah... | |
| 14:35:31 | efried | indeeed. | |
| 14:35:35 | jianghuaw_ | bauzas, I mean it automatically consumed display heads. but placement don't it. | |
| 14:35:51 | mdbooth | See, I actually think there's a lot of merit in having persistent driver state | |
| 14:36:14 | mdbooth | I think it should be local to the hypervisor, though | |
| 14:36:34 | efried | mdbooth We don't go crazy with persistent state, for sure. We save these slot mappings, and we save the partition's NVRAM. | |
| 14:36:55 | jianghuaw_ | bauzas, don't it => doesn't know of it. | |
| 14:37:16 | mdbooth | efried: Definitely in favour of the concept. | |
| 14:37:48 | efried | mdbooth I have to say, our implementation ain't pretty. | |
| 14:38:10 | jianghuaw_ | bauzas, virt driver can only stop a booting by checking the resoureces in the allocation. | |
| 14:38:21 | efried | mdbooth It was one of these rush jobs where we realized the problem (devices attached out of order) late in a cycle, and threw together a solution at the last minute. | |
| 14:38:55 | mdbooth | efried: Out of curiosity, why did you consider it a sufficiently big problem for a big rush job? | |
| 14:39:01 | jianghuaw_ | bauzas, or is there a way to correct the allocation if the allocated vGPU also supports display heads? | |
| 14:39:21 | mdbooth | Device ordering is also inconsistent in a Linux guest OS | |
| 14:39:29 | openstackgerrit | Chris Dent proposed openstack/nova master: [placement] re-use existing conf with auth token middleware https://review.openstack.org/523403 | |
| 14:39:38 | efried | mdbooth I don't remember all the reasons, but e.g. maybe you don't boot. | |
| 14:39:47 | efried | mdbooth Or possibly worse, boot to the wrong disk. | |
| 14:41:20 | efried | mdbooth Also could have something to do with applications in guests referring to devices by name - and if that name changes, you're kaput. | |
| 14:41:43 | mdbooth | efried: That last point isn't fixed by consistent device ordering | |
| 14:41:47 | efried | mdbooth Keep in mind we're running not just Linux, but also AIX and IBMi. | |
| 14:41:56 | mdbooth | Because as I say, Linux itself is also non-deterministic | |
| 14:42:31 | mdbooth | The first should be handled by the hypervisor specifying a boot order for its BIOS/UEFI | |
| 14:43:27 | mdbooth | Or whatever hopefully much better boot solution is in Power :) | |
| 14:43:43 | efried | mdbooth That's exactly the point: The boot list is specified based on hardware addresses, which are discovered based on slots. | |
| 14:45:04 | efried | mdbooth The big conceptual divide here is that PowerVM has a real hypervisor and the guests are real partitions. | |
| 14:46:10 | jianghuaw_ | bauzas, I think efried's opinion is that we shouldn't make inventory for VGPU_DISPLAY_HEAD as that's possible to be consumed automatically by requesting VGPU. | |
| 14:46:54 | efried | jianghuaw_ Only if display heads are *only* consumed automatically by requesting VGPU | |