Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-31
20:41:00 efried dansmith Well, yeah, what I'm trying to do is compose notes in enough detail - and with enough of the "easy" questions answered - that we can have a productive conversation in Denver (and not devolve into spoon metaphors).
20:41:09 edleafe artom: one of the things that came up in discussions in Austin about CAPI and such was that management of dynamic resources would never be handled by nova
20:41:15 dansmith but in reality, I think most cases will be satisfied by carving up your GPUs into chunks
20:41:34 edleafe so if your resources change, the virt driver has to report that change
20:41:49 artom edleafe, they would, I'm pretty sure
20:41:51 dansmith efried: sure, I know, but it'll go a lot quicker in person I think
20:42:06 artom edleafe, nova would still have to update its inventories though
20:42:37 edleafe artom: it does that periodically, based on what the virt driver tells it
20:42:40 efried Yeah, I'd like to understand more about how it breaks the model for me to change the inventory of something - for whatever reason I see fit.
20:42:45 artom edleafe, ah, we're fine then
20:42:51 efried ^ me == virt driver in this case
20:43:12 cristicalin hy, anyone have an idea why instances libvirt.xml may be missing metadata for user_id / project_id ?
20:43:13 dansmith efried: I said in response to allocation
20:43:14 edleafe efried: I always thought you looked like a virt driver
20:43:33 efried dansmith Well, yeah, it's sort of in response to allocation.
20:43:49 efried I guess I would call it in response to actually doing the virt driver work of satisfying the allocation.
20:43:51 dansmith efried: you should be exposing how much of a thing you have and the allocations come out of that
20:43:57 dansmith efried: you don't expose "available" you expose "total"
20:44:56 efried Hm, so I guess in this scenario I don't expose SRIOV_VF=48; instead I expose SRIOV_VF_EGRESS_PERCENTAGE_POINT=100
20:45:44 efried but that doesn't allow me to model the max #VFs as 48 at the same time. Or two separate VFs in the same claim.
20:46:13 efried ick
20:47:23 dansmith efried: regardless,
20:47:39 dansmith if your NIC has 10G of bandwidth, you can have your provider expose CUSTOM_GIGS=10,
20:47:55 dansmith and have your flavor consume some number of those to avoid over-subscribing
20:50:35 cfriesen_ efried: wouldn't you still say that you have 48 VFs total, and allocate from that?
20:50:56 openstackgerrit Matt Riedemann proposed openstack/nova master: Modernize set_vm_state_and_notify https://review.openstack.org/499799
20:51:07 dansmith presumably if you have 3 VFs left, but no GIGs left, you don't want to allocate any more
20:51:11 cfriesen_ efried: presumably you could also model total bandwidth, and divide it up nonuniformly between the VFs
20:52:16 efried dansmith So the VFs and the GIGs would be separate resource classes sitting next to each other, and you have to claim VF=1,GIG=2 or whatever in your request
20:52:28 efried cfriesen_ Could, but that kills the flexibility
20:52:45 efried cfriesen_ No idea at the outset how the end user is gonna want to split 'em up.
20:52:50 dansmith efried: yeah, so you end up with either one capping the things that can be there
20:53:13 dansmith either you run out of tiny VFs or GIGs from requests with large bandwidth requirements
20:53:19 tonygunk Just upgraded from Ocata to Pike and trying to upgrade nova DB - getting error: "AttributeError: 'module' object has no attribute 'get_rpc_transport'"
20:53:27 tonygunk http://paste.openstack.org/raw/620137/
20:53:28 efried dansmith So in that scenario a claim that specifies VF=1 but doesn't specify GIG would have to fail.
20:53:44 efried I can't default GIG because placement wouldn't know about it.
20:53:48 efried But
20:53:54 dansmith efried: that'd be up to the guy making out the flavors yeah
20:53:54 efried I would have to fail at compute
20:53:55 tonygunk Anyone know what is going on?
20:54:15 dansmith tonygunk: you have an outdated oslo_messaging I think
20:54:36 edleafe efried: sounds analogous to vcpu/ram/disk in splitting a physical server
20:54:46 edleafe efried: it's all in how the flavors are configured
20:54:48 dansmith edleafe: right, you waste some memory if you run out of disk
20:54:59 dansmith which is why we have flavors anyway,
20:55:07 dansmith so you can pack things so you don't waste stuff if you care about that
20:55:20 dansmith instead of just letting users request 1 cpu and 20TB of disk
20:55:25 tonygunk dansmith: Ok - thanks - I'll see if that resolves it
20:55:28 efried And a flavor that specifies CPU but not memory is useless.
20:55:51 efried Okay, I dig it. Thanks again.
20:56:23 cfriesen_ efried: I meant what dansmith is saying....two separate resource classes
20:56:52 efried (dansmith edleafe - that's the other thing I want to avoid in Denver - being the only one in the room needing to have basic concepts explained to him, and wasting everyone else's time.)
20:56:56 efried cfriesen_ Dig.
20:57:15 dansmith efried: understood, it's just not as fun for me if we do it this way
20:57:17 mriedem nova meeting in 3 minutes
20:57:55 cfriesen_ is it possible for an admin user to boot an instance "on behalf" of another user/project?
20:57:57 dansmith mriedem: dude I cannot effing wait
20:57:59 dansmith mriedem: PC BROS!
20:58:01 efried Sorry dansmith. I'll allocate you a predetermined number of units of your libation of choice.
20:58:15 dansmith efried: traits=wheat,unfiltered
20:58:25 efried noted
20:59:50 mriedem dansmith: better check your privilege
20:59:56 dansmith heh
21:01:49 tonygunk dansmith: now I'm getting this error after upgrading oslo.messaging to 5.30
21:01:52 tonygunk http://paste.openstack.org/raw/620138/
21:02:09 dansmith tonygunk: paste is old? :P
21:03:19 dansmith tonygunk: I can't help you as much with that one, sorry
21:11:49 tonygunk dansmith: actually uninstalling and reinstalling PasteDeploy was the solution.. weird..
21:12:03 dansmith *shrug* but good
21:12:28 efried tonygunk Seen that a hundred times. Curse PasteDeploy!
21:12:42 efried Sorry I didn't punch through your last paste, woulda been able to tell you right off.
21:13:03 tonygunk efried: :-)
21:22:38 efried I think if you give me nested resource providers, with some careful modeling I can do everything in the virt driver. The rest of nova doesn't need to know about a device as being any different from any other resource.
21:27:48 efried edleafe You mentioned earlier that "we still have some work to do with traits and nested providers before any of this will actually work" - somewhere I can find a summary of what's still on the table?
21:33:29 cfriesen_ efried: I think the PTG etherpad has some links
21:33:39 efried ah, cool
21:33:42 efried thanks cfriesen_
21:37:29 cfriesen efried: I think this is the nested providers work: https://review.openstack.org/#/q/topic:bp/nested-resource-providers,n,z
21:38:06 edleafe efried: yeah. There's the series from alex_xu starting with https://review.openstack.org/#/c/489206/, as well as Jay's nested RP series starting with https://review.openstack.org/#/c/470575/
21:38:17 efried edleafe cfriesen Thanks. Looks like not toooo much; hopefully containable in early q?
21:38:48 edleafe efried: only when you look at the code with light-red-tinted lenses
21:38:59 cfriesen efried: edleafe: I think there's also a missing piece to wire it up to live migration
21:39:02 efried Sigh.
21:39:17 efried I just thought you were blushing.
21:40:04 mriedem there are bugs that still need to be fixed from claims in the scheduler before we can add nested and shared RPs
21:40:16 mriedem the move operations are all over the floor
21:40:28 openstackgerrit Matt Riedemann proposed openstack/nova master: Modernize set_vm_state_and_notify https://review.openstack.org/499799
21:41:13 mriedem for anyone that cares, the main ones i'm tracking for backports to pike are
21:41:15 mriedem 1. https://bugs.launchpad.net/nova/+bug/1713786
21:41:16 openstack Launchpad bug 1713786 in OpenStack Compute (nova) "Allocations are not managed properly in all evacuate scenarios" [High,In progress] - Assigned to Matt Riedemann (mriedem)
21:41:26 mriedem 2. https://bugs.launchpad.net/nova/+bug/1713796
21:41:27 openstack Launchpad bug 1713796 in OpenStack Compute (nova) "Failed unshelve does not remove allocations from destination node" [High,Triaged]
21:41:47 mriedem i think that's it for right now
21:41:57 mriedem once those are fixed and backported to pike and merged, we can do a pike patch release
21:42:17 mriedem i've got the fix for half of the first bug up for review
21:42:27 bauzas roger.
21:42:28 mriedem starting here https://review.openstack.org/#/c/499678/2
21:43:08 mriedem tomorrow i'm going to work on the 2nd half, which is if evacuate fails on the compute, we have to cleanup the allocations created by the scheduler for the dest node
21:43:30 mriedem like, if the rebuild claim fials

Earlier   Later