| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2021-03-02 | |||
| 12:25:07 | lyarwood | yeah I'm not sure where some of the changes in the original hyperv change came from but pip was happy with these | |
| 12:25:47 | lyarwood | aaaaaand something has already failed | |
| 12:25:53 | sean-k-mooney | if we want to proceed with this for this cycle i wont try to block it but i do think we need to have a ptg discusssion about this and i dont think we should do this in the future | |
| 12:26:23 | sean-k-mooney | the new resolver has changed the meanin of LC and expanded its scope | |
| 12:26:47 | sean-k-mooney | if we want to continue to use that i think we need to discuss what LC now means | |
| 12:27:32 | lyarwood | sean-k-mooney: yeah I've added a line in the ptg pad but feel free to rephrase the question | |
| 12:29:06 | sean-k-mooney | lyarwood: run tox with -r to recreate the env | |
| 12:29:16 | sean-k-mooney | or before that check the pip version | |
| 12:29:37 | sean-k-mooney | it wont automatically upgrade the pip version | |
| 12:29:52 | sean-k-mooney | so if the enve was created with the old resolver then it will still be using it | |
| 12:30:39 | sean-k-mooney | oh | |
| 12:30:43 | lyarwood | that's a different job | |
| 12:30:50 | sean-k-mooney | lyarwood: you forgot to update requirements.txt | |
| 12:30:52 | lyarwood | I'm just about to run it locally now | |
| 12:31:00 | sean-k-mooney | you only bumped the min in lc | |
| 12:31:09 | lyarwood | for the in-direct deps? | |
| 12:31:13 | lyarwood | I thought that wasn't required? | |
| 12:31:17 | sean-k-mooney | it is | |
| 12:31:30 | sean-k-mooney | we also use those directly | |
| 12:31:30 | lyarwood | huh pip was fine without them | |
| 12:31:44 | sean-k-mooney | pip is but this is the requiremetns check job | |
| 12:32:10 | sean-k-mooney | we dont allow the min in requirements.txt to differ form lc by policy in openstack | |
| 12:32:47 | sean-k-mooney | and that job enforces that while also check the markers match what is in GR/UC | |
| 12:33:28 | lyarwood | *sigh* | |
| 12:33:29 | sean-k-mooney | lyarwood: the script that does this is in the requiremetns repo if i rememebr correctly its not in nova so you cant test this with tox in nova | |
| 12:33:42 | lyarwood | yeah just building the venv now | |
| 12:34:10 | sean-k-mooney | it wont see this error since this is not part of nova | |
| 12:34:54 | lyarwood | the requirements venv | |
| 12:34:59 | lyarwood | not nova | |
| 12:35:02 | sean-k-mooney | ah | |
| 12:35:28 | sean-k-mooney | anyway its a simple fix | |
| 12:35:37 | sean-k-mooney | just update nova's requiremetns.txt | |
| 12:35:44 | sean-k-mooney | with the same min version you set in lc | |
| 12:36:46 | sean-k-mooney | you do not need to change anything in the requiremetns repo in case that is not obvious form the error | |
| 12:37:46 | lyarwood | it's obvious, I just wanted to run the same test locally | |
| 12:37:53 | lyarwood | from the requirements repo | |
| 12:43:05 | ahsen | Hi, I'm getting an error while creating an instance. It says "Build of instance xxx aborted: Volume xxx did not finish being created even after we waited 188 seconds or 61 attempts. And its status is downloading." I did not get this error before and there is'nt any error on Cinder's logs. Do you have any idea why am I getting this error? How can I | |
| 12:43:06 | ahsen | increase waiting time or attemps? We are using Ussuri. Thank you. | |
| 12:46:36 | lyarwood | ahsen: the retries and interval are controlled by CONF.block_device_allocate_retries and CONF.block_device_allocate_retries_interval on the Nova side | |
| 12:46:50 | lyarwood | ahsen: but you should trace the request through to the Cinder side to understand why it's taking so long | |
| 12:52:09 | sean-k-mooney | ahsen: is it a large image or are you using HDDs on the cinder side or low bandwith nics | |
| 12:53:04 | sean-k-mooney | ahsen: i had to set block_device_allocate_retries_interval=10 on my home cluster | |
| 12:53:55 | sean-k-mooney | the time it took for qemu image to copy the image data for larger images was taking just over 60 seconds and it was timing out | |
| 12:54:17 | ahsen | lyarwood Thank you, I will try to increase those values. And actually Cinder creates volumes but I don't know how long does it take | |
| 12:55:10 | sean-k-mooney | ahsen: if you look at teh cidner driver log you should see the qemu-img command doing the data transfer | |
| 12:55:38 | sean-k-mooney | ahsen: for me it only became a proable for images over about 5-8Gs | |
| 12:56:53 | ahsen | sean-k-mooney Image is not large and we are using SSD also bandwith is not low | |
| 12:57:44 | lyarwood | ahsen: kk, you should see a request-id logged by Nova from Cinder that you can use to grep through your logs | |
| 12:57:54 | sean-k-mooney | ok increasing the interval to 10 might help but you should look at the cidner logs and try and determin why its taking so long in that case | |
| 12:58:12 | sean-k-mooney | ahsen: what cinder backend are you using by the way | |
| 12:58:38 | ahsen | sean-k-mooney I will look at them | |
| 12:58:51 | openstackgerrit | Lee Yarwood proposed openstack/nova master: requirements.txt: Bump os-brick to 4.2.0 https://review.opendev.org/c/openstack/nova/+/778177 | |
| 12:58:51 | openstackgerrit | Lee Yarwood proposed openstack/nova master: hyper-v rbd volume support https://review.opendev.org/c/openstack/nova/+/763550 | |
| 12:58:56 | lyarwood | okay lets try this again | |
| 13:00:28 | ahsen | lyarwood and sean-k-mooney I will try what you said. Thank you both | |
| 13:11:52 | openstackgerrit | Merged openstack/nova master: libvirt: add AsyncDeviceEventsHandler https://review.opendev.org/c/openstack/nova/+/772381 | |
| 13:26:44 | lyarwood | stephenfin: https://review.opendev.org/c/openstack/nova/+/673790/14/nova/virt/libvirt/host.py@1244 - going to grab some lunch but let me know if that concern isn't clear still. | |
| 13:43:34 | openstackgerrit | Elod Illes proposed openstack/nova stable/ussuri: Fallback to same-cell resize with qos ports https://review.opendev.org/c/openstack/nova/+/773932 | |
| 13:48:52 | gmann | brinzhang0: sorry i was away yesterday. replied on https://review.opendev.org/c/openstack/nova/+/766726/ | |
| 13:57:53 | bauzas | gibi: dansmith: huzzah \o/ the Compute RPC API bump patch eventually got a +1 from Zuul https://review.opendev.org/c/openstack/nova/+/761452 | |
| 14:03:28 | bauzas | fwiw, the grenade multinode job helped to find an issue | |
| 14:03:41 | bauzas | + some functest related to numa live migration | |
| 14:07:08 | gibi | bauzas: added the RPC bump to my review queue | |
| 14:11:51 | bauzas | gibi: I could split it by having two changes, one for the compute service manager and one for the compute client, but the existing change is quite simple to be reviewed | |
| 14:12:05 | gibi | thanks I will dig into it | |
| 14:15:47 | bauzas | gibi: I can help you to understand how this works | |
| 14:15:58 | sean-k-mooney | FYI we may need to rework how we do memory tracking to fix a previously unknow aspect of pci passhtough | |
| 14:16:16 | bauzas | gibi: once you begin to look at the RPC API, ping me and I'll explain | |
| 14:16:34 | sean-k-mooney | ill try and file a bug for it when i have time but basically memory oversubsciption cant be done if you have pci passthough/sriov | |
| 14:16:36 | bauzas | gibi: tl;dr: I'm providing a 5.x proxy for supporting old clients | |
| 14:16:44 | sean-k-mooney | it might also affect vgpu | |
| 14:16:59 | bauzas | sean-k-mooney: ack | |
| 14:17:06 | bauzas | sean-k-mooney: vgpu or gpu ? | |
| 14:17:10 | sean-k-mooney | both | |
| 14:17:17 | bauzas | why ? | |
| 14:17:18 | sean-k-mooney | we will have to verify it | |
| 14:17:33 | bauzas | b/c vgpu is different from pci passthrough and sriov | |
| 14:17:39 | sean-k-mooney | bauzas: libivrt is locking the guest memory pages whenever we use pci passthoug | |
| 14:17:50 | sean-k-mooney | it may or may not be doing the same for vgpu | |
| 14:17:55 | bauzas | ah | |
| 14:18:05 | bauzas | test it then, yes | |
| 14:18:39 | bauzas | and like I discussed with you, I'd also like to work on pci attach/detach | |
| 14:18:50 | sean-k-mooney | generic attach ya | |
| 14:19:22 | sean-k-mooney | this came up in a call i had this morining for vdpa when i was trying to dig into why it needed the memory to be locked explictly | |
| 14:19:43 | sean-k-mooney | for vdpa libvirt does not do it implcitly which is why i got dma error | |
| 14:20:02 | sean-k-mooney | so i have not actully avalidated it myself yet | |
| 14:20:22 | sean-k-mooney | once i do ill try and write up my findings | |
| 14:21:04 | sean-k-mooney | im not sure ill have time to try and validate vgpu but when i fiture out how to check if teh vm memory is locked maybe you can try? | |
| 14:28:15 | bauzas | sean-k-mooney: sure, I can try to get some vgpu environment (hopefully) | |
| 15:00:08 | sean-k-mooney | vgpus dont work with ubuntu hosts correct at leat not the nvida version? | |
| 15:00:50 | sean-k-mooney | actully im using teh mainline 5.11 kernel so never mind it wont install on my host anyway | |
| 15:01:06 | sean-k-mooney | the vdpa host im using also has a t4 but i dont really have a way to test it there | |
| 15:02:25 | bauzas | sean-k-mooney: no, ubuntu is not supported by nvidia drivers but you can use intel gvt-g if you really want to test vgpus | |
| 15:02:44 | bauzas | and a i915 host if you have one by hand | |
| 15:03:09 | sean-k-mooney | i was just wondering if the host i have would work or not. am maybe my laptop i could test with libvirt directly i guess | |
| 15:03:21 | bauzas | that's what I did for i915 | |
| 15:03:28 | bauzas | I just used an old laptop for testing | |
| 15:03:37 | sean-k-mooney | it wont happen in the next 2 weeks so we can figure it out later. | |
| 15:04:09 | bauzas | that's actually the problem with gvt-g, you can't find good intel cards that are for production | |
| 15:04:28 | bauzas | all the i915 cards are for desktop lines | |