| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-02 | |||
| 15:27:53 | mordred | efried: not this moment - I hit my nerd-snipe limit for the day - feel free to use/takeover/ignore that patch - or if I get stuck on other things later today and need a different mental task I may try poking again | |
| 15:28:23 | efried | mordred Rgr. My comments may be for myself, then :) | |
| 15:29:01 | mordred | efried: woot! | |
| 15:38:48 | mriedem | cdent: coming back to your question about the plan for fixing https://bugs.launchpad.net/nova/+bug/1707252 i don't know of one | |
| 15:38:48 | openstack | Launchpad bug 1707252 in OpenStack Compute (nova) "Claims in the scheduler does not account for doubling allocations on resize to same host" [Medium,Confirmed] | |
| 15:39:38 | mriedem | but it's not part of https://review.openstack.org/#/c/488510/ | |
| 15:39:47 | cdent | mriedem: yeah | |
| 15:40:28 | cdent | mriedem: the ideas I’ve heard batted around include: | |
| 15:40:40 | mriedem | the issue is here https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L196 | |
| 15:41:02 | cdent | adding migration uuids and adding the old allocation using the migration uuid, setting the consumer_uuid to the new allocation | |
| 15:41:25 | cdent | simply doubling in the move to same case and worrying about shared resource providers some other time (they complicate things rather a lot) | |
| 15:42:19 | mriedem | the migration uuid thing can't happen in pike | |
| 15:42:40 | cdent | I’m a bit confused on the real status of “shared providers”. Do you know? | |
| 15:43:00 | cdent | we can returned shared providers, but I’m not sure we create them | |
| 15:43:13 | cdent | if we don’t create them, then the simple doubling thing _might_ be made to work | |
| 15:43:14 | mriedem | we don't use them in any dsvm ci job | |
| 15:43:25 | mriedem | we don't create them in nova no | |
| 15:43:40 | mriedem | something like the ceph ci job would create a shared resource provider for disk | |
| 15:43:55 | mriedem | if we had a multinode ceph job, we could have 2 compute onde RPs and a single shared ceph disk RP | |
| 15:44:09 | cdent | yes, but we’d first need a resource tracker that was properly shared provider aware | |
| 15:44:13 | mriedem | that's on the etherpad for queens ptg to discuss | |
| 15:44:20 | mriedem | right, which we know it's not | |
| 15:44:29 | mriedem | which is https://bugs.launchpad.net/nova/+bug/1707256 | |
| 15:44:30 | cdent | that suggests, then, that we can proceed with a non-sharing solution for now | |
| 15:44:33 | openstack | Launchpad bug 1707256 in OpenStack Compute (nova) "Scheduler report client does not account for shared resource providers" [High,Confirmed] - Assigned to Jay Pipes (jaypipes) | |
| 15:45:02 | edleafe | So in a migration, would a shared resource, like disk, also be doubled? | |
| 15:45:24 | cdent | edleafe: long term, yes, but we don’t currently have them | |
| 15:45:25 | mriedem | it oculd | |
| 15:45:28 | mriedem | *could | |
| 15:45:43 | mriedem | if i'm resizing my disk from 20GB to 40GB, i need to account for that new disk allocatoin | |
| 15:45:56 | cdent | edleafe: and there’s also the cross case where a target host shares disk via two providers | |
| 15:46:12 | edleafe | mriedem: right, but would we account for the old 20GB and the new 40GB? Or just the larger? | |
| 15:46:17 | cdent | so you could migrate from: host a, disk x to host b, disk y OR host b, disk x | |
| 15:46:41 | edleafe | cdent: yeah, where the compute moves, but the disk doesn't | |
| 15:46:53 | cdent | edleafe: on the hangout earlier it was decided 60 | |
| 15:47:07 | cdent | that is, potentially over allocate, do the simple math | |
| 15:47:14 | edleafe | cdent: yeah, that's probably easier to implement | |
| 15:47:27 | cdent | in some situations it would be wrong, but not wrong broken | |
| 15:47:32 | edleafe | too many use cases overlapping | |
| 15:47:43 | cdent | quite | |
| 15:49:36 | cdent | mriedem, edleafe I could spike a doubling for resize to same host, but it wouldn’t be able to get started until about 4 hours from now | |
| 15:49:44 | dansmith | cdent: mriedem right I think for pike we have to double the allocation and then just gracefully subtract our old_flavor from the allocation if we're the only provider | |
| 15:49:54 | dansmith | and not worry about the shared stuff for the moment | |
| 15:51:05 | mriedem | there was a hangout earlier today? | |
| 15:51:15 | cdent | no earlier in the week | |
| 15:51:22 | mriedem | oh, there were a few :) | |
| 15:51:25 | cdent | (at least that’s what I was referring to) | |
| 15:53:33 | mriedem | so when we resize to the same host, this code is ending up with new_rp_uuids as empty, right? https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L200 | |
| 15:53:50 | mriedem | assuming no shared storage provider | |
| 15:54:13 | mriedem | which makes this False | |
| 15:54:13 | mriedem | if alloc['resource_provider']['uuid'] in new_rp_uuids: | |
| 15:54:24 | mriedem | so the dest rp uuid, which is the same as the source rp uuid, not get added in | |
| 15:54:27 | mriedem | for double the fun | |
| 15:54:40 | cdent | yeah, that’s my recollection | |
| 15:55:04 | mriedem | we could tell if this is a compute resource provider by checking for the VCPU resource class in the allocations | |
| 15:55:09 | mriedem | if that helps narrow things a bit | |
| 15:56:10 | cdent | mriedem: that’s one of the places we ended up on that aforementioned hangout | |
| 15:56:21 | cdent | but if we are discounting shared providers, for now, we don’t need to worry about that | |
| 15:56:46 | mriedem | doesn't seem like it would be hard to keep accounting for shared providers | |
| 15:56:53 | mriedem | if new_rp_uuids is empty, we know something is afoot | |
| 15:57:22 | mriedem | i think i need to check out gibi | |
| 15:57:30 | mriedem | gibi's new test for resize to same host first | |
| 15:57:39 | cdent | a) I’m thinkin in terms of trying to iterate, b) I think the code in jay’s stack is incorrect for shared providers anyway, based on some of the things dansmith said earlier in the week | |
| 15:57:40 | mriedem | and then could start playing with a change that builds on that | |
| 15:58:20 | dansmith | cdent: I assume so as well | |
| 15:58:38 | gibi | mriedem: that test needs a bit of update based on the above discussion. I used assert(max(old, new), usage) type of asserts but you agreed about old+new as I see | |
| 15:59:26 | cdent | gibi: you’re near the end of your day, yeah? | |
| 15:59:47 | gibi | cdent: yeah, and on a train with spotty conenction | |
| 16:00:03 | cdent | then I won’t say “maybe you should do the spike” :0 | |
| 16:00:19 | gibi | at least not today | |
| 16:00:27 | gibi | :) | |
| 16:00:50 | mriedem | so we know the scheduler report client doesn't know about shared storage providers, and will trample your shared storage provider allocations saying any disk consumed is local to the compute node provider | |
| 16:00:51 | dansmith | mriedem: we could probably help by landing his two bottom patches at least | |
| 16:01:01 | dansmith | mriedem: I +2d the bottom one this morning and can look at the next one now | |
| 16:01:15 | dansmith | mriedem: correct | |
| 16:01:33 | mriedem | so before we can say shared storage is supported, we have to fix that, and all of your computes have to be upgraded to the level that has that fix | |
| 16:01:44 | mriedem | meanwhile we're not doing a min service version check in the scheduler to account for that | |
| 16:02:19 | openstack | Launchpad bug 1707256 in OpenStack Compute (nova) "Scheduler report client does not account for shared resource providers" [High,Confirmed] - Assigned to Jay Pipes (jaypipes) | |
| 16:02:19 | bauzas | are folks discussing of https://bugs.launchpad.net/nova/+bug/1707256 ? | |
| 16:02:28 | mriedem | we're discussing all things | |
| 16:02:41 | bauzas | all things | |
| 16:02:55 | bauzas | :) | |
| 16:03:06 | mriedem | if we say, f it, shared storage isn't supported in pike, then we just fix the resize to same host thing by doubling allocations in the scheduler, right? | |
| 16:03:56 | dansmith | mriedem: yeah and ideally gracefully subtracting in the compute when done | |
| 16:04:13 | dansmith | mriedem: we should be as graceful as possible though so we don't screw up queens nodes that may do it right | |
| 16:04:16 | mriedem | on confirm resize? | |
| 16:04:19 | dansmith | yeah | |
| 16:04:43 | mriedem | hurts | |
| 16:04:43 | mriedem | ma | |
| 16:04:44 | mriedem | brain | |
| 16:06:58 | dansmith | this is oregon, | |
| 16:07:05 | dansmith | we have better things at our disposal | |
| 16:07:21 | cdent | my stash is cashed | |
| 16:08:24 | dansmith | mriedem: so I can start looking at making it do the right thing on the confirm if you want | |
| 16:09:10 | dansmith | I wish his top patch didn't marry the two things he's fixing together | |
| 16:09:18 | dansmith | the ocata compat and the resize_confirm fix | |
| 16:09:27 | dansmith | I'll put mine on top at least | |
| 16:10:46 | mriedem | i'll check out the bottom change that fixes PUT to overwrite all allocations - already did the other day and it made sense, seems simple, | |
| 16:11:00 | bauzas | I just +Wd it | |
| 16:11:05 | mriedem | also need to check out gibi's resize to same host tests, and then i was going to tinker with some of the code in https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L200 | |