| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-02 | |||
| 15:08:32 | stephenfin | I'm sure I'll survive in any case though :) | |
| 15:08:58 | ildikov | stephenfin: yeah, I didn't want to scare you with it, just wanted to be sure you're prepared :) | |
| 15:09:01 | mordred | efried: I just wrote most of the auth followup patch as a result of reading through to make sure I understood where the issues were ;) | |
| 15:09:08 | efried | mordred Figuring out how to get the auth into the glance client ... | |
| 15:09:15 | mordred | yup. done. patch coming | |
| 15:09:19 | sdague_ | mriedem: you misread my email about what's not being used in multinode | |
| 15:09:21 | stephenfin | ildikov: Ha, cheers :) | |
| 15:09:24 | efried | mordred Cool beans. | |
| 15:09:25 | mordred | (this is the nerd-sniping that just hapepned) | |
| 15:09:46 | mordred | efried: there is a test case that I didn't fix because it's a bit yuck and I don't know enough about the nova test cases | |
| 15:09:51 | mriedem | sdague_: i realized after i sent that | |
| 15:09:58 | ildikov | stephenfin: :) | |
| 15:10:00 | mordred | (basically, there's a test case that doesn't have any auth context info already prepared) | |
| 15:10:08 | mriedem | "it" with multiple nouns in the same sentence has been hurting me lately | |
| 15:10:31 | sdague_ | heh, no problem | |
| 15:10:35 | stephenfin | ildikov: I've learnt from my ill fated "Sure, you'll only need a light cotton sheet to sleep with" experience at Zurich last year. Turns out it gets colddd at night | |
| 15:10:51 | sdague_ | mriedem: I also responded, email should hit soon | |
| 15:11:05 | mriedem | i will prepare to respond to your response | |
| 15:12:18 | mriedem | wow, has anyone realized that a single trove change runs 26 jobs? | |
| 15:15:36 | ildikov | stephenfin: yeah, better to prepare as if you're lucky it gets down to 15 degrees Celsius at night the second half of next week | |
| 15:16:31 | ildikov | stephenfin: which I like way better than 27 even if I have to sleep in a tent, but we all have different tastes :) | |
| 15:22:38 | openstackgerrit | Monty Taylor proposed openstack/nova master: WIP Use auth from context for glance api servers https://review.openstack.org/490057 | |
| 15:22:41 | mordred | efried: ^^ | |
| 15:22:50 | efried | mordred Ack | |
| 15:23:01 | mordred | efried: I didn't fix nova/tests/unit/api/openstack/compute/test_images.py because of the auth context thing | |
| 15:23:13 | efried | mordred You gonna? | |
| 15:27:53 | mordred | efried: not this moment - I hit my nerd-snipe limit for the day - feel free to use/takeover/ignore that patch - or if I get stuck on other things later today and need a different mental task I may try poking again | |
| 15:28:23 | efried | mordred Rgr. My comments may be for myself, then :) | |
| 15:29:01 | mordred | efried: woot! | |
| 15:38:48 | openstack | Launchpad bug 1707252 in OpenStack Compute (nova) "Claims in the scheduler does not account for doubling allocations on resize to same host" [Medium,Confirmed] | |
| 15:38:48 | mriedem | cdent: coming back to your question about the plan for fixing https://bugs.launchpad.net/nova/+bug/1707252 i don't know of one | |
| 15:39:38 | mriedem | but it's not part of https://review.openstack.org/#/c/488510/ | |
| 15:39:47 | cdent | mriedem: yeah | |
| 15:40:28 | cdent | mriedem: the ideas I’ve heard batted around include: | |
| 15:40:40 | mriedem | the issue is here https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L196 | |
| 15:41:02 | cdent | adding migration uuids and adding the old allocation using the migration uuid, setting the consumer_uuid to the new allocation | |
| 15:41:25 | cdent | simply doubling in the move to same case and worrying about shared resource providers some other time (they complicate things rather a lot) | |
| 15:42:19 | mriedem | the migration uuid thing can't happen in pike | |
| 15:42:40 | cdent | I’m a bit confused on the real status of “shared providers”. Do you know? | |
| 15:43:00 | cdent | we can returned shared providers, but I’m not sure we create them | |
| 15:43:13 | cdent | if we don’t create them, then the simple doubling thing _might_ be made to work | |
| 15:43:14 | mriedem | we don't use them in any dsvm ci job | |
| 15:43:25 | mriedem | we don't create them in nova no | |
| 15:43:40 | mriedem | something like the ceph ci job would create a shared resource provider for disk | |
| 15:43:55 | mriedem | if we had a multinode ceph job, we could have 2 compute onde RPs and a single shared ceph disk RP | |
| 15:44:09 | cdent | yes, but we’d first need a resource tracker that was properly shared provider aware | |
| 15:44:13 | mriedem | that's on the etherpad for queens ptg to discuss | |
| 15:44:20 | mriedem | right, which we know it's not | |
| 15:44:29 | mriedem | which is https://bugs.launchpad.net/nova/+bug/1707256 | |
| 15:44:30 | cdent | that suggests, then, that we can proceed with a non-sharing solution for now | |
| 15:44:33 | openstack | Launchpad bug 1707256 in OpenStack Compute (nova) "Scheduler report client does not account for shared resource providers" [High,Confirmed] - Assigned to Jay Pipes (jaypipes) | |
| 15:45:02 | edleafe | So in a migration, would a shared resource, like disk, also be doubled? | |
| 15:45:24 | cdent | edleafe: long term, yes, but we don’t currently have them | |
| 15:45:25 | mriedem | it oculd | |
| 15:45:28 | mriedem | *could | |
| 15:45:43 | mriedem | if i'm resizing my disk from 20GB to 40GB, i need to account for that new disk allocatoin | |
| 15:45:56 | cdent | edleafe: and there’s also the cross case where a target host shares disk via two providers | |
| 15:46:12 | edleafe | mriedem: right, but would we account for the old 20GB and the new 40GB? Or just the larger? | |
| 15:46:17 | cdent | so you could migrate from: host a, disk x to host b, disk y OR host b, disk x | |
| 15:46:41 | edleafe | cdent: yeah, where the compute moves, but the disk doesn't | |
| 15:46:53 | cdent | edleafe: on the hangout earlier it was decided 60 | |
| 15:47:07 | cdent | that is, potentially over allocate, do the simple math | |
| 15:47:14 | edleafe | cdent: yeah, that's probably easier to implement | |
| 15:47:27 | cdent | in some situations it would be wrong, but not wrong broken | |
| 15:47:32 | edleafe | too many use cases overlapping | |
| 15:47:43 | cdent | quite | |
| 15:49:36 | cdent | mriedem, edleafe I could spike a doubling for resize to same host, but it wouldn’t be able to get started until about 4 hours from now | |
| 15:49:44 | dansmith | cdent: mriedem right I think for pike we have to double the allocation and then just gracefully subtract our old_flavor from the allocation if we're the only provider | |
| 15:49:54 | dansmith | and not worry about the shared stuff for the moment | |
| 15:51:05 | mriedem | there was a hangout earlier today? | |
| 15:51:15 | cdent | no earlier in the week | |
| 15:51:22 | mriedem | oh, there were a few :) | |
| 15:51:25 | cdent | (at least that’s what I was referring to) | |
| 15:53:33 | mriedem | so when we resize to the same host, this code is ending up with new_rp_uuids as empty, right? https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L200 | |
| 15:53:50 | mriedem | assuming no shared storage provider | |
| 15:54:13 | mriedem | if alloc['resource_provider']['uuid'] in new_rp_uuids: | |
| 15:54:13 | mriedem | which makes this False | |
| 15:54:24 | mriedem | so the dest rp uuid, which is the same as the source rp uuid, not get added in | |
| 15:54:27 | mriedem | for double the fun | |
| 15:54:40 | cdent | yeah, that’s my recollection | |
| 15:55:04 | mriedem | we could tell if this is a compute resource provider by checking for the VCPU resource class in the allocations | |
| 15:55:09 | mriedem | if that helps narrow things a bit | |
| 15:56:10 | cdent | mriedem: that’s one of the places we ended up on that aforementioned hangout | |
| 15:56:21 | cdent | but if we are discounting shared providers, for now, we don’t need to worry about that | |
| 15:56:46 | mriedem | doesn't seem like it would be hard to keep accounting for shared providers | |
| 15:56:53 | mriedem | if new_rp_uuids is empty, we know something is afoot | |
| 15:57:22 | mriedem | i think i need to check out gibi | |
| 15:57:30 | mriedem | gibi's new test for resize to same host first | |
| 15:57:39 | cdent | a) I’m thinkin in terms of trying to iterate, b) I think the code in jay’s stack is incorrect for shared providers anyway, based on some of the things dansmith said earlier in the week | |
| 15:57:40 | mriedem | and then could start playing with a change that builds on that | |
| 15:58:20 | dansmith | cdent: I assume so as well | |
| 15:58:38 | gibi | mriedem: that test needs a bit of update based on the above discussion. I used assert(max(old, new), usage) type of asserts but you agreed about old+new as I see | |
| 15:59:26 | cdent | gibi: you’re near the end of your day, yeah? | |
| 15:59:47 | gibi | cdent: yeah, and on a train with spotty conenction | |
| 16:00:03 | cdent | then I won’t say “maybe you should do the spike” :0 | |
| 16:00:19 | gibi | at least not today | |
| 16:00:27 | gibi | :) | |
| 16:00:50 | mriedem | so we know the scheduler report client doesn't know about shared storage providers, and will trample your shared storage provider allocations saying any disk consumed is local to the compute node provider | |
| 16:00:51 | dansmith | mriedem: we could probably help by landing his two bottom patches at least | |
| 16:01:01 | dansmith | mriedem: I +2d the bottom one this morning and can look at the next one now | |