Earlier  
Posted Nick Remark
#openstack-cyborg - 2017-05-17
15:15:42 zhipeng not necessary live-attachment
15:15:51 zhipeng but just attach.detach ops in general
15:16:00 zhipeng for VM instances
15:16:15 zhipeng since you will need the instance id and host id anyway
15:16:22 zhipeng and nova got all these
15:16:26 jkilpatr ok then, we have a booted vm we need to do the attachment, we can setup nova to do the passthrough and then reboot the instance, not sure how nova feels about it but if it doesn't work I don't think they would be opposed to us making that work
15:16:37 jkilpatr but also nova does support live pci attachment, so we can/should just use that
15:17:07 zhipeng yes if Nova indeed could support that
15:17:14 jkilpatr https://wiki.openstack.org/wiki/Nova/pci_hotplug
15:17:25 jkilpatr this is a stub right now, so some support
15:17:28 zhipeng we just don't differentiate in Cyborg
15:18:14 jkilpatr the driver actually handles this of course, we just run attach on a live instance that was spawned with a flavor or some other indicator that says "cyborg: TeslaP100"
15:18:59 zhipeng yes
15:19:21 jkilpatr because we have to make sure Nova gets it to the right spot firs, so we need to have the placement api fed live knowlege and then the instance needs to call the resource name we fed the placement api
15:19:40 jkilpatr maybe we can have a command in cyborg that will help you make flavors, but that's for later.
15:19:51 zhipeng hotplug could be the trait when we use placement api
15:20:33 zhipeng so when we schedule it, it know it needs a compute node with hotplug feature
15:20:46 jkilpatr wouldn't we want a bunch of traits
15:20:58 zhipeng yep I guess so :)
15:21:03 jkilpatr like gpu, with cuda support, with hotplug support .... so on and so forth
15:21:10 zhipeng we could start with really basic and simple ones
15:21:11 jkilpatr this is why we need a flavor creation wizard in cyborg but once again, later.
15:21:25 zhipeng yes
15:21:34 jkilpatr ok anyone have comments on these ideas? things that they might want that this won't cover?
15:22:15 jkilpatr zhipeng, I'm going to put a comment on your api patch as a reminder to add cinder like attach/detach to the spec sound good?
15:23:00 zhipeng which I thought is already done in the current patch ?
15:23:27 cdent it would be great to see, at some point, a narration of the expected end to end flow from a user's standpoint, if it doesn't already exist. Including how various services will be touched.
15:24:17 zhipeng cdent we got a flow chart in our BOS presentation, but rather rudimentary at the moment
15:24:20 jkilpatr cdent, we have a decent idea of how we want it to work but I expect some things will change as we get into the nitty gritty of placement problems
15:24:42 cdent sure, change is the nature of this stuff :)
15:24:47 jkilpatr zhipeng, ok so if I want to attach an accelerator to an instance what do I do? Do I put to update an accelerator spec with a new instance ID to attach to?
15:24:48 zhipeng :)
15:25:18 jkilpatr cdent, I think I'll put up a user workflow spec later today, just so that we keep track of all of this better.
15:25:30 cdent \o/
15:25:48 cdent can you add me as a review on that when it is up, so I get some email to remind me to look?
15:26:23 zhipeng just as we drew for our presentation, after the user using Cyborg service to complete the discovery phase and Cyborg finishing interaction with placement to advertise the accelerator inventory
15:26:24 jkilpatr will do, whats your email?
15:26:44 cdent cdent@anticdent.org is me on gerrit
15:27:01 zhipeng then user just request to create an instance on a compute node with the corresponding accelerator trait
15:27:32 zhipeng if trait include hotplug, then maybe it will be a live attachment
15:27:45 jkilpatr I really don't think we're going to get away with one trait per accelerator, users will probably bundle them into flavors, but instead of being tied to a list of whitelisted pci devices these flavors can be much mroe general.
15:27:50 zhipeng which means user could attach the accelerator after VM creation
15:28:35 zhipeng I was told by jay today that trait are per resource provider
15:29:02 zhipeng so it would mostly be one trait per compute node
15:29:08 jkilpatr I'll have to look at it in detail, I was watching the summit presentation again today.
15:29:17 zhipeng or we got vGPUs or FPGA virtual functions
15:29:28 zhipeng then it would be nested resource provider
15:29:38 zhipeng and we could have trait on the virtual functions
15:29:53 zhipeng but anyways it does not tie to a specific accelerator
15:30:10 zhipeng just depending on you model your accelerators into resource providers
15:30:22 zhipeng cdent plz correct me if i'm wrong :P
15:30:24 jkilpatr we're going to need to be careful with that.
15:31:03 zhipeng #link https://pbs.twimg.com/media/DAAUxEWUAAAV6zA.jpg
15:31:16 cdent zhipeng: that looks mostly correct, but I'm only partially paying attention :(
15:31:27 zhipeng cdent no problemo :)
15:31:51 zhipeng as long as I don't make any extremely wrong claims :)
15:32:07 jkilpatr so implementation. What can we start and when?
15:32:24 zhipeng as soon as we freeze the specs
15:32:35 zhipeng i suppose we should all go ahead start coding
15:32:48 crushil Can we focus on closing out the specs out first though?
15:32:55 zhipeng yes
15:32:57 jkilpatr ok then that's a plan.
15:33:15 zhipeng okey, then for api spec
15:33:31 zhipeng #link https://review.openstack.org/445814
15:33:35 zhipeng any other questions ?
15:34:22 jkilpatr I just posted a comment there, otherwise I'm happy enough
15:34:31 zhipeng okey
15:34:54 jkilpatr um should we pick a database tech? what's available already sql, mongo... reddis (not sure about that one)
15:35:08 crushil MariaDB?
15:35:15 zhipeng #action jkilpatr to post a reminder comment, the api spec patch is LGTM
15:35:39 zhipeng i think we could just use mysql
15:35:54 jkilpatr MariaDB == mysql except when it doesnt
15:36:04 jkilpatr I think openstack ships with Maria right now
15:36:07 zhipeng yes
15:36:09 crushil yup
15:36:28 zhipeng next up, agent spec
15:36:33 zhipeng #link https://review.openstack.org/#/c/446091/
15:37:12 jkilpatr looks like most people are happy with it.
15:37:14 zhipeng #info Jay Pipes suggest agent could directly interact with placement api, instead of going through conductor
15:37:30 zhipeng #link https://pbs.twimg.com/media/DAAUtyoUMAAXQXI.jpg
15:37:50 jkilpatr so from the summit preso all the computes already talk to the placement api themselves
15:37:51 zhipeng but I guess we don't need to reflect that in the agent spec
15:37:54 jkilpatr so it's designed to scale well like that
15:38:30 jkilpatr I'd prefer to be explicit, I'll patch it into my spec today
15:38:48 crushil We should reflect that in the agent spec
15:39:01 zhipeng jkilpatr yes, and for implementation, Jay suggest we could directly just copy nova/scheduler/client/report.py
15:39:24 zhipeng since it is basically rest calls between agent and placement api
15:39:32 jkilpatr that's the sort of laziness I can get behind.
15:39:42 zhipeng XD
15:39:44 crushil lol
15:40:29 zhipeng #agreed jkilpatr do a quick update on agent spec to reflect jaypipes comment, then the agent spec patch LGTM
15:40:38 zhipeng okey, next up, generic driver
15:40:58 zhipeng #link https://review.openstack.org/#/c/447257/
15:41:04 zhipeng any more comments
15:41:08 zhipeng looks fine to me
15:41:52 crushil jkilpatr, Any other comments on your end? I have tried to address all of your and Roman's comments in the patch
15:41:56 jkilpatr what about detect accelerator? discovery has to be handled by someone, do we want drivers to have a discovery call?
15:42:10 jkilpatr I like the rest of the api list for it, good job
15:43:12 crushil I can add that to the list. What would be the flow though for discovery?
15:43:37 zhipeng i think discovery already part of the spec ?

Earlier   Later