| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-05-12 | |||
| 09:00:48 | sean-k-mooney | but what you are trying to do is going against the grain of how nova is intended to be used ot a degree | |
| 09:00:52 | bauzas | gibi: ack thanks | |
| 09:03:05 | sean-k-mooney | jkulik: im not sure how configurable the vmware dirver is | |
| 09:03:30 | sean-k-mooney | but do you want to reserve space for this giant vms or no? | |
| 09:04:13 | openstack | Launchpad bug 1838309 in OpenStack Compute (nova) "Live migration might fail when run after revert of previous live migration" [Undecided,Won't fix] - Assigned to Vladyslav Drok (vdrok) | |
| 09:04:13 | bauzas | gibi: agreed with Wontfix https://bugs.launchpad.net/nova/+bug/1838309 ? | |
| 09:04:57 | sean-k-mooney | what you could do woudl be to reserve ram on the specifci hosts where you intend to spawn these isntances. then create a custom resouce class for them and a flavor that request 2TB or ram but set resocues:memory_mb=0 resouces:CUSTOM_GIANT_VM=1 | |
| 09:05:00 | jkulik | sean-k-mooney, not in general. we currently run something, that frees up enough space on one note - so we can spawn one. | |
| 09:05:32 | sean-k-mooney | ok so that wont work if you want it to be dynamic then | |
| 09:05:33 | jkulik | yeah, we have something like that. but it has to fit into normal quota for billing | |
| 09:05:47 | tony_su | gibi: This is Tony Su and I am doing provider-config-file re-propose thing. Could you kindly spare some time to review its spec (only minor changes in History and Assignee sections vs. Ussuri version). https://review.opendev.org/#/c/725788/ | |
| 09:06:22 | sean-k-mooney | jkulik: well the flavor.ram value would still be 2TB | |
| 09:06:32 | jkulik | sean-k-mooney, we dynamically create a sharing child-resource-provider for a node, that has the CUSTOM_GIANT_VM resource | |
| 09:07:23 | sean-k-mooney | ok but the issue is you still need to aling the flavor.ram to what that means for that host | |
| 09:07:24 | jkulik | sean-k-mooney, so currently, it works for us. but the flavor's RAM doesn't align with NUMA on all nodes. meaning, we would have to create more flavors. | |
| 09:07:31 | sean-k-mooney | based on the hardware version | |
| 09:07:35 | jkulik | yes | |
| 09:08:12 | sean-k-mooney | jkulik: may i ask why you actully use VM instead of ironic nodes at that scale | |
| 09:08:19 | jkulik | but given that we only free up one node, that could fit 2 TB either because it's 2 TB in size or because it's bigger, the 2 TB are more a rounded value the custerom might request | |
| 09:08:32 | tony_su | gibi: Really appreciate your support and help. | |
| 09:08:58 | jkulik | sean-k-mooney, sure. we were on ironic, but our customers would like to have the automatic failover when a node goes down, that vmware provides. also TCO. | |
| 09:10:25 | sean-k-mooney | i see so they are relying on the plathform for failover rahter then using an orchestration layer like k8s to mage there application | |
| 09:10:39 | ikla | the controller and node doesn't need to have the same device in in for passthrough? | |
| 09:10:40 | sean-k-mooney | so this is very much a pets not cattel usecase | |
| 09:10:50 | jkulik | since the customer can't find out, which precise flavor she can currently deploy via API, it would have been nice to have specific sub-flavors for ~ 2 TB | |
| 09:10:52 | gibi | bauzas, stephenfin: easy +2 https://review.opendev.org/#/c/721278 (ussuri spec move) | |
| 09:10:57 | ikla | my flavor keeps failing when trying to turn on a instance | |
| 09:11:03 | sean-k-mooney | ikla: the alias need to be on the computes and contoler | |
| 09:11:14 | jkulik | sean-k-mooney, yes. definitely pets | |
| 09:11:18 | sean-k-mooney | ikla: the pci whitelist is only needed on the compute | |
| 09:11:27 | ikla | alias yes, device only on compute.. | |
| 09:11:28 | ikla | k | |
| 09:11:52 | sean-k-mooney | ikla: so the contole just uses it to translate the value in the flavor into pci request for schduling | |
| 09:12:20 | openstack | Launchpad bug 1838309 in OpenStack Compute (nova) "Live migration might fail when run after revert of previous live migration" [Undecided,Won't fix] - Assigned to Vladyslav Drok (vdrok) | |
| 09:12:20 | gibi | bauzas: I agree with WontFix https://bugs.launchpad.net/nova/+bug/1838309 | |
| 09:12:29 | bauzas | cool, moving on | |
| 09:13:03 | gibi | tony_su: thanks for taking that feature over. I will look at the spec soon | |
| 09:15:27 | openstack | Launchpad bug 1878024 in OpenStack Compute (nova) "disk usage of the nova image cache is not counted as used disk space" [Undecided,Confirmed] | |
| 09:15:27 | gibi | bauzas: thanks for looking at the new bugs. You have the bug lock, I have yet another downstream originated issue to look at https://bugs.launchpad.net/nova/+bug/1878024 | |
| 09:16:12 | sean-k-mooney | gibi: thats a tricky one | |
| 09:16:40 | sean-k-mooney | in that really we should never fail to boot a vm becasue of a cached image taking up space | |
| 09:17:10 | sean-k-mooney | i.e. i would expect as an operator for the cache to be perurged first | |
| 09:17:25 | sean-k-mooney | so im not sure i would want it to be counted as used | |
| 09:17:47 | sean-k-mooney | that siad we dont purge the cache today as far as i know at least not while the image is used on the host | |
| 09:17:54 | sean-k-mooney | so i can see this causeing issues too | |
| 09:18:18 | sean-k-mooney | could you just workaround this using the host reserved disk paramater | |
| 09:18:47 | sean-k-mooney | or perhaps we need a config option to limit the size of the cache | |
| 09:19:37 | gibi | sean-k-mooney: the cache is used as a backing file for qcow root fs so it cannot be really purged | |
| 09:19:40 | sean-k-mooney | the easy option is just to add the size of the cache to the reserved value in placment but that will reduce the amount of vms that can be spawned | |
| 09:19:54 | sean-k-mooney | gibi: i tought we had two copies | |
| 09:20:01 | sean-k-mooney | on in the cache and a second for the backing file | |
| 09:20:20 | sean-k-mooney | isnt the cach module indpentend of the image backend | |
| 09:20:22 | gibi | sean-k-mooney: reserved_host_disk_mb would only be useful if we could limit the size of the cache | |
| 09:21:02 | sean-k-mooney | yep the two line "fix" is to jsut count the cache size and added it to the reserved value in placment | |
| 09:21:10 | sean-k-mooney | but i dont think that is a resonable long term fix | |
| 09:22:45 | gibi | simply counting it as reserved does not help when a VM is booted with a new image to a compute. There we would need to make sure that both the new cached image and the VM root fs fits the compute | |
| 09:22:58 | gibi | we discussed these options with dansmith yesterday | |
| 09:23:09 | gibi | see the summary in the bug and the link to the IRC discussion | |
| 09:24:14 | openstackgerrit | Merged openstack/nova-specs master: move implemented specs in ussuri https://review.opendev.org/721278 | |
| 09:25:35 | sean-k-mooney | gibi: are we sure we use the cache image directly for the backing file by the way | |
| 09:26:13 | sean-k-mooney | gibi: i tought that was how it worked but i remeber someone telling me its not how it worked and that we chage two copies one in the caceh and one for the backing file | |
| 09:26:59 | gibi | let me double check | |
| 09:27:10 | ikla | any info on gpu passthrough w/ rtx 8000's ? | |
| 09:27:27 | huaqiang | stephenfin: I am not sure how to keep exisiting +2, will that be kept if I do not change the | |
| 09:27:38 | sean-k-mooney | ikla: are you trying to do full passthough or are you trying to use its vgpu capablity | |
| 09:27:46 | aarents | I think it is still the same directory | |
| 09:28:14 | huaqiang | will the +2 be kepet if I donot change the 'Change-ID'? | |
| 09:28:53 | gibi | sean-k-mooney: http://paste.openstack.org/show/793424/ | |
| 09:29:35 | gibi | sean-k-mooney: based on this the image in the _base dir is the backing store for the instance's image | |
| 09:29:58 | gibi | I have two instances from the same image, and have one backing image in the _base dir | |
| 09:30:22 | sean-k-mooney | gibi: yep but is that the cached image or is it a second copy | |
| 09:30:41 | sean-k-mooney | depenidn gon how we set the force cow and force raw values this might change | |
| 09:31:03 | sean-k-mooney | in that case we are using a raw backing file | |
| 09:31:15 | sean-k-mooney | but waht formate was the image originally? | |
| 09:31:26 | sean-k-mooney | i assume qcow? | |
| 09:31:27 | gibi | sean-k-mooney: a) it doesn't matter. I can reword the bug that the bacing file size is not counted as used b) I think this is the cache as if I delete the instances it does not deleted imediately | |
| 09:31:59 | bauzas | ikla: not sure I understand your question about gpu passthrough | |
| 09:33:19 | gibi | sean-k-mooney: I will try to look into the case when the VM uses a raw file to see if then the image in _base is deleted, or if I can make a small change to be deleted | |
| 09:34:46 | sean-k-mooney | gibi: there is a periodic that deletes the image when no vms is using it on the host as far as i am aware | |
| 09:35:02 | gibi | sean-k-mooney: yes, that is the cache manager :) | |
| 09:35:05 | sean-k-mooney | so until that runs the cached copy will be there | |
| 09:35:48 | gibi | sean-k-mooney: anyhow the core of the problem is that nova uses more disk that it counts as used and today there is no way to avoid that except putting the _base dir on a separate partition | |
| 09:36:46 | gibi | when we had the DiskFilter it had the disk_available_least information from the compute to prevent overallocating the disks but placement does not have such info | |
| 09:37:19 | sean-k-mooney | well not entirely | |
| 09:37:32 | ikla | grid vgpu | |
| 09:37:46 | sean-k-mooney | the disk avaiable least behavior was still affected by the disk allocation ratio | |
| 09:38:26 | bauzas | ikla: gpu passthrough != grid vgpu , you know ? | |
| 09:38:36 | sean-k-mooney | gibi: i still go back to https://gist.github.com/JCallicoat/43505cab0535057ca4fb every time i want to figure that out | |
| 09:39:03 | bauzas | in one case, you're litterally giving up the gpu device to the guest | |
| 09:39:18 | ikla | sorry I didn't explain that correctly | |
| 09:39:26 | ikla | you could do a full passthrough | |
| 09:39:27 | bauzas | in the other case, you're asking the nvidia driver to slice your gpu into pieces that can be provided to the guests | |
| 09:39:28 | ikla | or do vgpu | |
| 09:39:34 | ikla | correct | |
| 09:39:42 | sean-k-mooney | ikla: yes both are supportred | |
| 09:39:53 | bauzas | so, again, what's your question ? | |
| 09:40:00 | ikla | is there docs on it | |
| 09:40:09 | sean-k-mooney | you can use mdev based vgpus or you can do direct passhtough of the gpu | |
| 09:40:14 | bauzas | ikla: indeed | |
| 09:40:24 | bauzas | https://docs.openstack.org/nova/latest/admin/virtual-gpu.html | |