| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-10-05 | |||
| 20:45:19 | jaypipes | edleafe: are you still with us? | |
| 20:46:03 | edleafe | I had to walk away for a while. I was getting way too frustrated that this thing that we have said all along was an opaque blob now has versioned information | |
| 20:47:07 | edleafe | No one seems to consider that an a-r is a placement artifact, not a scheduler/conductor thing | |
| 20:47:36 | edleafe | it is created by placement, and is used only by placment, and is never persisted | |
| 20:47:42 | edleafe | it is always "the latest" | |
| 20:48:04 | edleafe | But we want to treat it like it is a versioned data structure that scheduler "can know about" | |
| 20:49:21 | cdent | (what ed just said about “always latest” is part of why I’m agnostic on the client side. It would be fundamentally correct for the report client to send the header as ‘latest’ because the placement service is always its own latest) | |
| 20:50:03 | dansmith | cdent: but the client has to construct the actual request | |
| 20:50:20 | dansmith | right now it puts the user/project in there, and may have to do other things later | |
| 20:50:33 | dansmith | so it can't say latest for the request, it has to say a version | |
| 20:50:42 | cdent | yes, that’s where the notion of opaque blob falls apart | |
| 20:50:53 | edleafe | dansmith: right now it just passes back the a-r to claim. It doesn't build anything | |
| 20:50:55 | dansmith | the other thing is, | |
| 20:51:03 | edleafe | the user/project is in the a-r | |
| 20:51:07 | dansmith | the scheduler cannot communicate with placement at version latest | |
| 20:51:14 | dansmith | and it does have to look at the results of those calls | |
| 20:51:14 | jaypipes | edleafe: not currently it isn't, no. | |
| 20:51:40 | jaypipes | edleafe: the scheduler client's claim_resources() method adds user id and project id to the HTTP request payload. | |
| 20:51:41 | dansmith | so the compute or conductor saying latest _cannot_ be right forever | |
| 20:51:46 | edleafe | jaypipes: https://github.com/openstack/nova/blob/master/nova/api/openstack/placement/handlers/allocation.py#L77 | |
| 20:52:25 | openstackgerrit | Dan Smith proposed openstack/nova master: Revert allocations by migration uuid https://review.openstack.org/498949 | |
| 20:52:25 | openstackgerrit | Dan Smith proposed openstack/nova master: Pre-create migration object https://review.openstack.org/498950 | |
| 20:52:26 | openstackgerrit | Dan Smith proposed openstack/nova master: Make migration uuid hold allocations for migrating instances https://review.openstack.org/506420 | |
| 20:52:26 | openstackgerrit | Dan Smith proposed openstack/nova master: Refactor resource tracker to account for migration allocations https://review.openstack.org/506419 | |
| 20:52:27 | openstackgerrit | Dan Smith proposed openstack/nova master: Make live migration hold resources with a migration allocation https://review.openstack.org/507638 | |
| 20:52:34 | jaypipes | edleafe: that's the request payload for PUT /allocations, not the format of the allocation_request object that is returned in the GET /allocation_candidates HTTP response. | |
| 20:53:24 | jaypipes | edleafe: https://github.com/openstack/nova/blob/master/nova/api/openstack/placement/handlers/allocation_candidate.py#L57-L66 | |
| 20:53:53 | jaypipes | edleafe: we don't currently add the user and project ID into the allocation_request object in the return from GET /allocation_candidates | |
| 20:54:04 | jaypipes | edleafe: unfortunately. was an oversight on my part. | |
| 20:54:18 | jaypipes | edleafe: I'm sure dansmith at some point told me to put it in there and I just forgot. | |
| 20:54:36 | dansmith | it doesn't matter, because 'latest' is not the version used by the scheduler (in the future when we're doing things correctly and placement is external) | |
| 20:56:24 | cdent | so if scheduler and placement are out of sync, grind | |
| 20:56:56 | dansmith | when placement is external, we must tolerate them being out of sync | |
| 20:59:32 | edleafe | oh, geez, I give up. I missed this: https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L154 | |
| 20:59:48 | edleafe | Forget everything I said about a-rs being opaque | |
| 21:00:00 | mriedem | yes i remember pointing out late in pike that we needed to update nova-status' check for the required minimum placement microversion to be 1.10 because that's what the scheduler was requesting during claim_resources | |
| 21:00:01 | edleafe | that ship has sailed. | |
| 21:00:14 | mriedem | we really only needed 1.8 for the user_id/project_id thing (i think?) | |
| 21:00:56 | mriedem | and we had to put something in the release notes saying you have to make sure to upgrade placement before scheduler since scheduler requires this new higher microversoin in placement that wasn't available in ocata | |
| 21:01:08 | edleafe | I'm going to finish the stuff I've been trying to work on and then I'll rethink how to change the series to add a versioned allocation_request to the Selection object | |
| 21:01:29 | mriedem | if the pike scheduler was requesting 'latest' to an ocata placement, the request might pass at whatever 'latest' is for placement in ocata, but not what the pike scheduler client actually needs | |
| 21:02:02 | dansmith | edleafe: okay and you caught the bit I said about the selectionlist object potentially being okay if we're going to use it for holding a version right? | |
| 21:03:04 | edleafe | dansmith: yeah, but that's minor | |
| 21:03:48 | dansmith | edleafe: yep, just saying, if you wan to go back to doing it that way, I'm cool with it | |
| 21:03:50 | cdent | edleafe’s link raises another wart doesn’t it? If _move_operation_alloc_request is working in the guts of alloc request, it has to know the version | |
| 21:04:21 | cdent | is that called from only the scheduler, or also in the cells? | |
| 21:04:28 | cdent | (and presumably there are others like it?) | |
| 21:04:38 | jaypipes | cdent: we're trying to get rid of that entirely. | |
| 21:04:45 | jaypipes | cdent: and do the migration owns allocation thing. | |
| 21:04:51 | edleafe | cdent: it will be called from within the cells too | |
| 21:04:52 | mriedem | cdent: it's called from the scheduler and, for the time being, superconductor | |
| 21:04:55 | cdent | yes, but will still inspect don’t we? | |
| 21:04:58 | mriedem | during force live migrate and force evacuate | |
| 21:05:01 | mriedem | where the scheduler is skipped | |
| 21:05:06 | mriedem | edleafe: not within the cells | |
| 21:05:18 | cdent | and in any case that code is pike | |
| 21:05:21 | mriedem | edleafe: oh you mean with alternate hosts yeah | |
| 21:05:22 | edleafe | mriedem: the cell conductor will have to claim | |
| 21:05:26 | mriedem | right right | |
| 21:05:42 | dansmith | but we don't need too look inside the a-r in the claim during reschedule | |
| 21:05:52 | mriedem | we just proxy it through | |
| 21:05:55 | dansmith | this move claim thing is a good example of the scheduler needing to examine it | |
| 21:05:57 | dansmith | mriedem: right | |
| 21:06:02 | mriedem | "here is a request the scheduler told me to send at this version k?!" | |
| 21:06:09 | mriedem | "<3 cell conductor" | |
| 21:06:26 | edleafe | dansmith: and I was thinking that this move claim thing is a bad example | |
| 21:06:48 | dansmith | edleafe: it's a bad example in the cosmic sense of things that suck... yes :) | |
| 21:11:04 | efried | mriedem Please cast your critical eye upon https://review.openstack.org/#/c/488137/ when you get time. | |
| 21:11:38 | efried | IIRC the goal was to get the whole series in by the first milestone thingy. | |
| 21:11:48 | mriedem | efried: oh efried | |
| 21:11:57 | mriedem | did i actually say that was a goal? | |
| 21:12:05 | mriedem | i mentioned it as being doable during a meeting a few weeks back | |
| 21:12:12 | mriedem | and you've been cruising my house at 1am ever since | |
| 21:12:15 | efried | Oh, please let me find that. Stand by. | |
| 21:13:19 | efried | mriedem http://eavesdrop.openstack.org/meetings/nova/2017/nova.2017-09-21-14.00.log.txt @14:03:09 | |
| 21:14:00 | efried | It's entirely possible I've been misinterpreting "...should ... have ... merged by then" | |
| 21:14:11 | mriedem | that can be interpreted so many different ways | |
| 21:14:15 | mriedem | would never hold up in a court | |
| 21:14:41 | jaypipes | can we finalize on a decision on this then? | |
| 21:14:53 | efried | Fair enough. But if it ever gets interpreted as "this should have merged by then," I don't want it to be because I didn't pester people for reviews :) | |
| 21:16:24 | mriedem | efried: don't worry i know you've asked several times | |
| 21:16:36 | mriedem | jaypipes: pass the version down | |
| 21:17:14 | dansmith | yep | |
| 21:17:16 | edleafe | jaypipes: [t 4Bt] | |
| 21:17:17 | purplerbot | <edleafe> I'm going to finish the stuff I've been trying to work on and then I'll rethink how to change the series to add a versioned allocation_request to the Selection object [2017-10-05 21:01:08.583741] [n 4Bt] | |
| 21:17:36 | jaypipes | ok, thanks edleafe | |
| 21:26:45 | mriedem | ha, | |
| 21:26:49 | mriedem | good news folks | |
| 21:26:57 | mriedem | get ready for 100% nova gate failure | |
| 21:28:21 | dansmith | wat | |
| 21:28:28 | dansmith | I see the top of our gate is failing | |
| 21:28:33 | mriedem | yeah i know what it is | |
| 21:28:41 | mriedem | but i'll never tell | |
| 21:28:49 | dansmith | just fix, I don't care if you tell | |
| 21:29:01 | mtreinish | mriedem: oh it's your new test | |
| 21:29:10 | mriedem | SHHHHHHHHHHHHHHHH | |
| 21:29:14 | mriedem | TREINISH! | |
| 21:29:29 | mriedem | who were the ad wizards that merged that one | |
| 21:29:47 | mtreinish | mriedem: oomichi_afk gave it the +W | |
| 21:29:49 | dansmith | the shelve offload one? | |