| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-cyborg - 2018-03-15 | |||
| 16:21:06 | openstackgerrit | Nguyen Van Trung proposed openstack/cyborg master: Add default configuration files to data_files https://review.openstack.org/553463 | |
| 20:51:18 | openstackgerrit | Li Liu proposed openstack/cyborg master: Implemented all the APIs specified in cyborg-fpga-model-proposal.rst Added table Attribute Added Objects physical_function and virtual_function Added foreign key constrain for deployable to accelerator Added an API deployable_get_by_filters Added corespon https://review.openstack.org/552734 | |
| #openstack-cyborg - 2018-03-16 | |||
| 02:00:48 | openstackgerrit | Li Liu proposed openstack/cyborg master: Implemented all the APIs specified in cyborg-fpga-model-proposal.rst https://review.openstack.org/552734 | |
| 03:33:07 | openstackgerrit | Li Liu proposed openstack/cyborg master: Implemented the Objects and APIs for vf/pf https://review.openstack.org/552734 | |
| 03:58:37 | openstackgerrit | Li Liu proposed openstack/cyborg master: Implemented the Objects and APIs for vf/pf https://review.openstack.org/552734 | |
| 05:10:13 | openstackgerrit | Li Liu proposed openstack/cyborg master: Implemented the Objects and APIs for vf/pf https://review.openstack.org/552734 | |
| 05:37:02 | openstackgerrit | Li Liu proposed openstack/cyborg master: Implemented the Objects and APIs for vf/pf https://review.openstack.org/552734 | |
| 05:47:40 | Sundar | Hi Howard, if you are here, can we chat briefly? | |
| 06:05:27 | openstackgerrit | Li Liu proposed openstack/cyborg master: Implemented the Objects and APIs for vf/pf https://review.openstack.org/552734 | |
| 13:03:04 | kosamara | Hi! I tried getting cyborg on my devstack, but 'stack.sh' fails to start up the cyborg services. They complain that: | |
| 13:03:13 | kosamara | Exception: Versioning for this project requires either an sdist tarball, or access to an ups┤ | |
| 13:03:15 | kosamara | tream git repository. It's also possible that there is a mismatch between the package name in setup.cfg and the argument given to ┤ | |
| 13:03:17 | kosamara | pbr.version.VersionInfo. Project name cyborg was given, but was not able to be found.y | |
| 13:04:41 | kosamara | I've included the cyborg plugin in devstack's "local.conf" like this: | |
| 13:04:50 | kosamara | "enable_plugin cyborg git://git.openstack.org/openstack/cyborg stable/queens" | |
| 13:05:30 | kosamara | Does anyone have any suggestion on the cause and what to do? Thanks! | |
| 13:22:53 | zhipengh[m] | Hi kosamara , there is a bug with devstack and a patch is on the way to address that | |
| 13:33:13 | kosamara | thanks zhipengh! | |
| #openstack-cyborg - 2018-03-19 | |||
| 15:23:39 | openstackgerrit | Li Liu proposed openstack/cyborg master: Implemented the Objects and APIs for vf/pf https://review.openstack.org/552734 | |
| 19:24:25 | openstackgerrit | Li Liu proposed openstack/cyborg master: Implemented the Objects and APIs for vf/pf https://review.openstack.org/552734 | |
| #openstack-cyborg - 2018-03-20 | |||
| 21:54:29 | openstackgerrit | Sundar Nadathur proposed openstack/cyborg master: Specification for Cyborg/Nova interaction for scheduling. https://review.openstack.org/554717 | |
| #openstack-cyborg - 2018-03-21 | |||
| 03:07:51 | openstackgerrit | YumengBao proposed openstack/cyborg-specs master: Create an initial cyborg-specs repository using cookiecutter https://review.openstack.org/554766 | |
| 14:00:17 | zhipeng | #startmeeting openstack-cyborg | |
| 14:00:21 | openstack | Meeting started Wed Mar 21 14:00:17 2018 UTC and is due to finish in 60 minutes. The chair is zhipeng. Information about MeetBot at http://wiki.debian.org/MeetBot. | |
| 14:00:22 | openstack | Useful Commands: #action #agreed #help #info #idea #link #topic #startvote. | |
| 14:00:25 | openstack | The meeting name has been set to 'openstack_cyborg' | |
| 14:00:41 | zhipeng | #topic Roll Call | |
| 14:00:49 | zhipeng | #info Howard | |
| 14:01:11 | shaohe_feng_ | hi zhipeng | |
| 14:01:18 | kosamara | hi! | |
| 14:01:23 | zhipeng | hi guys | |
| 14:01:33 | crushil | \o | |
| 14:01:35 | zhipeng | please use #info to record you names | |
| 14:01:48 | kosamara | #info Konstantinos | |
| 14:01:50 | Sundar | #info Sundar | |
| 14:02:15 | crushil | #info Rushil | |
| 14:02:57 | shaohe_feng_ | #info shaohe | |
| 14:05:02 | Sundar | Do we have a Zoom link? | |
| 14:05:06 | zhipeng | let's wait for a few more minutes in case more people will join | |
| 14:05:14 | zhipeng | Sundar no we only have irc today | |
| 14:05:23 | Sundar | Sure, Zhipeng. Thanks | |
| 14:06:18 | Yumeng__ | #info Yumeng__ | |
| 14:09:45 | zhipeng | okey let's start | |
| 14:09:57 | zhipeng | #topic CERN GPU use case introduction | |
| 14:10:08 | zhipeng | kosamara plz takes us away | |
| 14:10:41 | kosamara | I joined you recently, so let me remind that I'm a technical student at CERN, integrating GPUs into our openstack. | |
| 14:10:56 | kosamara | Our use case is computation only at this point. | |
| 14:11:31 | kosamara | We have implemented and are currently testing a nova-only pci-passthrough. | |
| 14:12:18 | kosamara | We intend to also explore vGPUs, but they don't seem to fit our use case very much: licensing costs, limited CUDA support (nvidia case both). | |
| 14:12:47 | kosamara | At the moment the big issues are enforcing quotas on GPUs and security concerns. | |
| 14:13:07 | kosamara | 1. Subsequent users can potentially access data on the GPU memory | |
| 14:13:39 | kosamara | 2. Low-level access means they could change the firmware or even cause the host to restart | |
| 14:14:23 | kosamara | We are looking at a way to mitigate at least the first issue, by performing some kind of cleanup on the GPU after use | |
| 14:14:55 | kosamara | But in our current workflow this would require a change in nova. | |
| 14:15:13 | kosamara | It looks like something that could be done in cyborg? | |
| 14:15:42 | zhipeng | i think so, i remember we had similar convo regarding clean up on FPGA in the PTG | |
| 14:15:43 | Li_Liu | are you try to force a "reset" after usage of the device? | |
| 14:16:02 | kosamara | Reset of the device? | |
| 14:16:32 | kosamara | According to a research article, in nvidia's case performing a device reset through nvidia-smi is not enough for the data leaks. | |
| 14:16:42 | Li_Liu | something like that. like Zhipeng said, "clean up" | |
| 14:16:49 | kosamara | - article: https://www.semanticscholar.org/paper/Confidentiality-Issues-on-a-GPU-in-a-Virtualized-E-Maurice-Neumann/693a8b56a9e961052702ff088131eb553e88d9ae | |
| 14:17:19 | kosamara | The additional complexity in pci passthrough is that the host can't access the GPUs (no drivers) | |
| 14:17:39 | shaohe_feng_ | so once the GPU devices are detached, cyborg should do clean up at once. | |
| 14:17:40 | Sundar | What kind of clean up do you have in mind? Zero out the RAM? | |
| 14:18:15 | kosamara | My thinking is to put the relevant resources in a special state after deallocation from their previous VM and use a service VM to perform the cleanup ops themselves. | |
| 14:18:21 | Li_Liu | memset() to zero for all ? | |
| 14:18:27 | kosamara | Yes, zero the ram | |
| 14:19:00 | kosamara | But ideally we would like to ensure the firmware's state is valid too. | |
| 14:19:13 | shaohe_feng_ | only this one method to clear up? | |
| 14:19:47 | kosamara | Sorry, I'm out of context. This method is from where? | |
| 14:19:58 | shaohe_feng_ | zero the ram | |
| 14:20:15 | kosamara | Yes for the ram. | |
| 14:20:30 | shaohe_feng_ | OK, got it. | |
| 14:20:35 | kosamara | But we would also like to check the firmware's state. | |
| 14:20:54 | kosamara | And at least ensure that it hasn't been tampered with. | |
| 14:21:22 | shaohe_feng_ | so no ways to power off the GPU device, and re-power it? | |
| 14:21:23 | zhipeng | i think this is doable with cyborg | |
| 14:21:43 | kosamara | Not that I know of, apart from rebooting the host. | |
| 14:21:44 | zhipeng | shaohe_feng_ that will mess up the state ? | |
| 14:21:59 | Sundar | The service VM needs to run in every compute node, and it needs to have all the right drivers for various GPU devices. We need to see how practical that is. | |
| 14:22:14 | Li_Liu | we might need to get some help from Nvidia | |
| 14:22:48 | zhipeng | let's say we have the NVIDIA driver support | |
| 14:22:50 | Sundar | In a vGPU scenario, how do you power off the device without affecting other users? | |
| 14:22:52 | zhipeng | for cyborg | |
| 14:22:59 | kosamara | Sundar: unless vfio is capable of doing these basic operations? | |
| 14:23:39 | zhipeng | do we still need the service VM ? | |
| 14:24:11 | kosamara | I haven't explored practically nvidia's vGPU scenario. The host is supposed to be able to operate on the GPU in that case. | |
| 14:24:32 | Sundar | kosamara: Not sure I understand. vfio is a generic module for allowing device access. How would it know about specific nvidia devices? | |
| 14:24:44 | shaohe_feng_ | kosamara: do you use vfio pci passthrough or pci-stub passthrough? | |
| 14:25:08 | kosamara | zhipeng: vfio-pci is the host stub driver when the gpu is passed through. If we could do something through that, then we wouldn't need the service VM. | |
| 14:25:32 | kosamara | shaohe_feng_: vfio-pci | |
| 14:25:37 | zhipeng | kosamara good to know | |
| 14:26:16 | kosamara | Sundar: yes, I'm just making a hypothesis. I expect it can't, but perhaps zeroing-out the ram is a general enough operation and it can do it...? | |
| 14:27:40 | Vipparthy | hi | |
| 14:27:49 | shaohe_feng_ | kosamara: is zeroing-out the ram time consuming? | |
| 14:28:20 | Sundar | kosamara: I am not a Nvidia expert. But presumably, to access the RAM, one would need to look at the PCI BAR space and find the base address or something? If so, that would be device-specific | |
| 14:29:49 | kosamara | Sundar: thanks for the input. I can research the feasibility of this option this week. | |
| 14:30:15 | zhipeng | kosamara I will also talk to the NVIDIA OpenStack team about it | |
| 14:30:26 | zhipeng | see if we could come out with something for Rocky :) | |
| 14:30:45 | kosamara | shaohe_feng_: I don't have a number right now. It should depend on the manner: if it happens on the device without pci transports it should be quite fast. | |