Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-02
15:39:47 cdent mriedem: yeah
15:40:28 cdent mriedem: the ideas I’ve heard batted around include:
15:40:40 mriedem the issue is here https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L196
15:41:02 cdent adding migration uuids and adding the old allocation using the migration uuid, setting the consumer_uuid to the new allocation
15:41:25 cdent simply doubling in the move to same case and worrying about shared resource providers some other time (they complicate things rather a lot)
15:42:19 mriedem the migration uuid thing can't happen in pike
15:42:40 cdent I’m a bit confused on the real status of “shared providers”. Do you know?
15:43:00 cdent we can returned shared providers, but I’m not sure we create them
15:43:13 cdent if we don’t create them, then the simple doubling thing _might_ be made to work
15:43:14 mriedem we don't use them in any dsvm ci job
15:43:25 mriedem we don't create them in nova no
15:43:40 mriedem something like the ceph ci job would create a shared resource provider for disk
15:43:55 mriedem if we had a multinode ceph job, we could have 2 compute onde RPs and a single shared ceph disk RP
15:44:09 cdent yes, but we’d first need a resource tracker that was properly shared provider aware
15:44:13 mriedem that's on the etherpad for queens ptg to discuss
15:44:20 mriedem right, which we know it's not
15:44:29 mriedem which is https://bugs.launchpad.net/nova/+bug/1707256
15:44:30 cdent that suggests, then, that we can proceed with a non-sharing solution for now
15:44:33 openstack Launchpad bug 1707256 in OpenStack Compute (nova) "Scheduler report client does not account for shared resource providers" [High,Confirmed] - Assigned to Jay Pipes (jaypipes)
15:45:02 edleafe So in a migration, would a shared resource, like disk, also be doubled?
15:45:24 cdent edleafe: long term, yes, but we don’t currently have them
15:45:25 mriedem it oculd
15:45:28 mriedem *could
15:45:43 mriedem if i'm resizing my disk from 20GB to 40GB, i need to account for that new disk allocatoin
15:45:56 cdent edleafe: and there’s also the cross case where a target host shares disk via two providers
15:46:12 edleafe mriedem: right, but would we account for the old 20GB and the new 40GB? Or just the larger?
15:46:17 cdent so you could migrate from: host a, disk x to host b, disk y OR host b, disk x
15:46:41 edleafe cdent: yeah, where the compute moves, but the disk doesn't
15:46:53 cdent edleafe: on the hangout earlier it was decided 60
15:47:07 cdent that is, potentially over allocate, do the simple math
15:47:14 edleafe cdent: yeah, that's probably easier to implement
15:47:27 cdent in some situations it would be wrong, but not wrong broken
15:47:32 edleafe too many use cases overlapping
15:47:43 cdent quite
15:49:36 cdent mriedem, edleafe I could spike a doubling for resize to same host, but it wouldn’t be able to get started until about 4 hours from now
15:49:44 dansmith cdent: mriedem right I think for pike we have to double the allocation and then just gracefully subtract our old_flavor from the allocation if we're the only provider
15:49:54 dansmith and not worry about the shared stuff for the moment
15:51:05 mriedem there was a hangout earlier today?
15:51:15 cdent no earlier in the week
15:51:22 mriedem oh, there were a few :)
15:51:25 cdent (at least that’s what I was referring to)
15:53:33 mriedem so when we resize to the same host, this code is ending up with new_rp_uuids as empty, right? https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L200
15:53:50 mriedem assuming no shared storage provider
15:54:13 mriedem if alloc['resource_provider']['uuid'] in new_rp_uuids:
15:54:13 mriedem which makes this False
15:54:24 mriedem so the dest rp uuid, which is the same as the source rp uuid, not get added in
15:54:27 mriedem for double the fun
15:54:40 cdent yeah, that’s my recollection
15:55:04 mriedem we could tell if this is a compute resource provider by checking for the VCPU resource class in the allocations
15:55:09 mriedem if that helps narrow things a bit
15:56:10 cdent mriedem: that’s one of the places we ended up on that aforementioned hangout
15:56:21 cdent but if we are discounting shared providers, for now, we don’t need to worry about that
15:56:46 mriedem doesn't seem like it would be hard to keep accounting for shared providers
15:56:53 mriedem if new_rp_uuids is empty, we know something is afoot
15:57:22 mriedem i think i need to check out gibi
15:57:30 mriedem gibi's new test for resize to same host first
15:57:39 cdent a) I’m thinkin in terms of trying to iterate, b) I think the code in jay’s stack is incorrect for shared providers anyway, based on some of the things dansmith said earlier in the week
15:57:40 mriedem and then could start playing with a change that builds on that
15:58:20 dansmith cdent: I assume so as well
15:58:38 gibi mriedem: that test needs a bit of update based on the above discussion. I used assert(max(old, new), usage) type of asserts but you agreed about old+new as I see
15:59:26 cdent gibi: you’re near the end of your day, yeah?
15:59:47 gibi cdent: yeah, and on a train with spotty conenction
16:00:03 cdent then I won’t say “maybe you should do the spike” :0
16:00:19 gibi at least not today
16:00:27 gibi :)
16:00:50 mriedem so we know the scheduler report client doesn't know about shared storage providers, and will trample your shared storage provider allocations saying any disk consumed is local to the compute node provider
16:00:51 dansmith mriedem: we could probably help by landing his two bottom patches at least
16:01:01 dansmith mriedem: I +2d the bottom one this morning and can look at the next one now
16:01:15 dansmith mriedem: correct
16:01:33 mriedem so before we can say shared storage is supported, we have to fix that, and all of your computes have to be upgraded to the level that has that fix
16:01:44 mriedem meanwhile we're not doing a min service version check in the scheduler to account for that
16:02:19 bauzas are folks discussing of https://bugs.launchpad.net/nova/+bug/1707256 ?
16:02:19 openstack Launchpad bug 1707256 in OpenStack Compute (nova) "Scheduler report client does not account for shared resource providers" [High,Confirmed] - Assigned to Jay Pipes (jaypipes)
16:02:28 mriedem we're discussing all things
16:02:41 bauzas all things
16:02:55 bauzas :)
16:03:06 mriedem if we say, f it, shared storage isn't supported in pike, then we just fix the resize to same host thing by doubling allocations in the scheduler, right?
16:03:56 dansmith mriedem: yeah and ideally gracefully subtracting in the compute when done
16:04:13 dansmith mriedem: we should be as graceful as possible though so we don't screw up queens nodes that may do it right
16:04:16 mriedem on confirm resize?
16:04:19 dansmith yeah
16:04:43 mriedem ma
16:04:43 mriedem hurts
16:04:44 mriedem brain
16:06:58 dansmith this is oregon,
16:07:05 dansmith we have better things at our disposal
16:07:21 cdent my stash is cashed
16:08:24 dansmith mriedem: so I can start looking at making it do the right thing on the confirm if you want
16:09:10 dansmith I wish his top patch didn't marry the two things he's fixing together
16:09:18 dansmith the ocata compat and the resize_confirm fix
16:09:27 dansmith I'll put mine on top at least
16:10:46 mriedem i'll check out the bottom change that fixes PUT to overwrite all allocations - already did the other day and it made sense, seems simple,
16:11:00 bauzas I just +Wd it
16:11:05 mriedem also need to check out gibi's resize to same host tests, and then i was going to tinker with some of the code in https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L200
16:12:24 cfriesen mriedem: would we be looking at backporting any shared storage accounting fix back to pike? or only fixing in Q? (or does that sort of depend on what the fix looks like?)
16:13:06 mriedem it depends, but also not sure since the scheduler isn't making any distinction based on the version of the compute serivce
16:13:07 mriedem *Service
16:13:22 mriedem so it kind of sucks to say, 'well make sure you have this fix and everything is upgraded first'
16:13:45 cdent dansmith: are you doing just the undoubling, or also the doubling?
16:13:58 dansmith cdent: we're already doubling right?

Earlier   Later