Earlier  
Posted Nick Remark
#openstack-nova - 2018-09-14
15:32:54 bauzas of course, it would require a reboot
15:33:42 pvc_ the nvidia grid will work on tesla?
15:34:04 pvc_ here https://www.nvidia.co.uk/download/driverResults.aspx/99148/en-uk
15:35:04 bauzas https://docs.nvidia.com/grid/6.0/grid-vgpu-user-guide/index.html#install-vgpu-package-generic-linux-kvm are the KVM related notes
15:37:02 pvc_ thank you so much :) i just need to download the latest nvidia grid driver and thats is
15:37:03 pvc_ :)
15:38:10 bauzas pvc_: and a P100 is said to be supported with nvidia grid driver 6.2 https://docs.nvidia.com/grid/6.0/grid-vgpu-user-guide/index.html#vgpu-types-tesla-p100-12gb
15:39:54 pvc_ thank you so much. i will just first where to install the nvidia grid driver 6.2
15:40:21 bauzas out of curiosity, which nvidia driver version were you using ?
15:40:25 bauzas was it a grid one ?
15:40:37 bauzas (the nvidia licensing is so confusing...)
15:41:42 bauzas mriedem: efried: FWIW, I tested booting a server within the functional test and it graciously fails
15:41:48 bauzas which proves our point
15:42:16 bauzas now, I'll upload that one-liner WIP and test it against the nested-allocations series
15:42:17 mriedem bauzas: cool. can you rebase tetsuro's fixes on top and see it pass?
15:42:19 efried bauzas: How about when you add the use-nested-allocation-candidates code to it?
15:42:23 efried ... yeah, that :)
15:42:36 bauzas yeah I'm on it
15:42:53 pvc_ for now i dont have nvidia-driver i will find an nvidia driver for linux because when i see the grid it only supporte citrix vsphare bauzas
15:42:56 bauzas tetsuro: fine with me rebasing ?
15:43:04 tetsuro bauzas: sure
15:43:12 bauzas tetsuro: thanks
15:43:31 bauzas pvc_: nvidia grid drivers are licensed you know
15:43:41 bauzas you can't just pull one from internet :)
15:44:13 pvc_ i see so i cant use this: https://www.nvidia.com/Download/driverResults.aspx/136954/en-us for my p100 tesla
15:44:23 bauzas nope
15:44:30 pvc_ i need to contact nvidia for the licensing
15:44:33 bauzas but nvidia provides some evaluation drivers for testing
15:44:41 bauzas so you can start playing
15:45:04 pvc_ noted on this i will find that :)
15:45:14 pvc_ here https://www.nvidia.com/download/driverResults.aspx/129776/en-us?
15:45:21 pvc_ but for vmware only
15:45:28 bauzas nope, there http://www.nvidia.com/object/grid-evaluation.html
15:46:15 pvc_ thank you so much for your effort and time. im so sorry for disturbing you :)
15:46:45 pvc_ 90 day trial :)
15:47:34 bauzas pvc_: you don't disturb me, no worries :)
15:48:07 pvc_ have a great day. thank you so much. :)
15:48:10 mriedem bauzas: i wonder if we should have some kind of known issues section of the vgpu doc for stuff like this?
15:48:20 mriedem i.e. for things people run into
15:48:29 openstackgerrit Jay Pipes proposed openstack/nova-specs master: Standardize CPU resource tracking https://review.openstack.org/555081
15:48:52 mriedem bauzas: like we know you have to mask the hypervisor for nvidia right?
15:48:52 bauzas mriedem: we have an open bug for vgpu docs
15:49:10 bauzas but I'm torn on providing vendor specific documentation
15:49:16 mriedem i'm not,
15:49:18 mriedem let's add it
15:49:27 mriedem they don't have to be super detailed,
15:49:28 bauzas ok so there is a bug for it
15:49:34 mriedem just "remember to check this""
15:49:38 bauzas https://bugs.launchpad.net/nova/+bug/1752463
15:49:39 openstack Launchpad bug 1752463 in OpenStack Compute (nova) "Attaching virtual GPU devices to guests in nova" [Medium,Confirmed]
15:49:52 bauzas we can amend the existing docs to be a bit more oriented
15:49:52 mriedem pvc_: can you summarize your issue in that bug ^ ?
15:49:54 mriedem so we can document it?
15:50:07 mriedem i think it's just that your trial period expired?
15:50:24 bauzas mriedem: pvc_'s issue was that he wasn't using the nvidia grid driver
15:50:34 pvc_ ahm no. this is my first time using Tesla p100
15:50:43 pvc_ im using the gtx 1030 before just the gpu passthrough
15:51:32 pvc_ i just having a problem on getting the available vgpus for my card then bauzas taught me all ;)
15:52:26 pvc_ since i cannot list the directory that is here in the document https://docs.openstack.org/nova/queens/admin/virtual-gpu.html
15:53:27 bauzas running vgpus is like aligning the planets
15:53:44 mriedem if we can point to vendor docs or something i think that's ok
15:54:05 bauzas you need both a "certified" server, a recent kernel, some nvidia product line and some proprietary very specific driver
15:54:28 bauzas yeah, I think you're right
15:54:32 lyarwood kashyap: https://review.openstack.org/#/c/602477/1 - if you have time, I've found a weird block-job-complete error from QEMU v2.11.0, have you seen this before?
15:54:45 bauzas we can somehow add some vendor information without being too much specific
15:54:45 kashyap lyarwood: Hey
15:54:54 kashyap lyarwood: Let me quickly look
15:54:56 pvc_ thank you for time mriedem and bauzas. god bless :)
15:55:06 kashyap lyarwood: The job completion area is hairy :-(
15:55:43 kashyap lyarwood: See the glorious discussion (we also had this on the 'openstack-dev' list some moons ago): https://bugzilla.redhat.com/show_bug.cgi?id=1382165
15:55:43 openstack bugzilla.redhat.com bug 1382165 in libvirt "virDomainGetBlockJobInfo: Adjust job reporting based on QEMU stats & the "ready" field of `query-block-jobs`" [Unspecified,New] - Assigned to libvirt-maint
15:57:15 openstackgerrit Jay Pipes proposed openstack/nova-specs master: Standardize CPU resource tracking https://review.openstack.org/555081
15:58:42 kashyap lyarwood: I'd be interested to see if you hit that twice or more in a row
15:59:45 lyarwood kashyap: kk, I'll try to check later
16:12:31 kashyap lyarwood: So, apparently this is "expected with QEMU 2.11". Checking w/ the QEMU block dev guy who fiddled with it
16:12:46 kashyap I'll write a comment on the Gerrit change. You can process it async.
16:12:53 lyarwood kashyap: ack thanks
16:19:52 kashyap lyarwood: If you see in the log, several ('drive-mirror') jobs _were_ completed sucessfully
16:20:22 mriedem TOOT! https://review.openstack.org/#/c/588665/
16:20:25 bauzas mriedem: bug me with changes if you want reviews
16:20:27 bauzas hah
16:22:07 kashyap lyarwood: Okay, I can confirm from log examination that this is another racy occurence of the libvirt bug I pointed out
16:22:33 lyarwood kashyap: cool, should we work around this for now?
16:23:21 kashyap lyarwood: There is no "one" simple workaround :-( We need to make Nova consume the 'events' from libvirt / QEMU block layer
16:23:35 kashyap The problem is here. If you see this line:
16:23:36 kashyap 2018-09-13 22:57:16.517+0000: 30124: info : qemuMonitorJSONIOProcessLine:213 : QEMU_MONITOR_RECV_REPLY: mon=0x7f67a0021980 reply={"return": [{"io-status": "ok", "device": "drive-virtio-disk1", "busy": true, "len": 1073741824, "offset": 1073741824, "paused": false, "speed": 0, "ready": false, "type": "mirror"}], "id": "libvirt-143"}
16:23:42 kashyap You see "len" == "offset"
16:23:48 kashyap Which can mean, okay, swap has completed.
16:24:11 kashyap _But_: You don't have the "ready": *true* <-- That is the critical missing bit for the job to complete.
16:26:56 kashyap Anyway.
16:29:32 lyarwood kashyap: kk thanks, about to talk about your libvirt/qemu bump on L774 https://etherpad.openstack.org/p/nova-ptg-stein btw
16:30:29 kashyap lyarwood: Yeah, that's fairly mundane and boring stuff that needs to be done
16:35:05 mriedem sorrison: https://review.openstack.org/#/c/591607/
16:35:12 bauzas tetsuro: I'm confused by your branching on nested-alloc-candidates
16:36:06 bauzas tetsuro: https://review.openstack.org/#/c/585672/ seems to be the top change, but you have https://review.openstack.org/#/c/527728/ that requires a rebase on the last version of the above ?
16:36:07 tetsuro bauzas: Is that about https://review.openstack.org/#/c/585672/4?
16:36:21 bauzas yup
16:36:29 bauzas ok, you know why ?
16:36:57 bauzas I'll put my WIP on top of the last rev of 585672
16:37:07 bauzas and then we'll figure out later

Earlier   Later