| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-02 | |||
| 18:56:50 | mriedem | dansmith: aren't we trying to fix the computes to not stomp on things the scheduler is doing? | |
| 18:57:13 | mriedem | like, make the compute aware of shared storage allocations | |
| 18:57:13 | dansmith | mriedem: not really | |
| 18:57:20 | dansmith | not ocata computes | |
| 18:57:26 | mriedem | right, pike computes | |
| 18:57:35 | jaypipes | mriedem: no. we're trying to fix pikes to not stomp on things that other pike computes may have been doing. | |
| 18:57:41 | dansmith | right | |
| 18:57:52 | dansmith | while still being compatible with ocata computes | |
| 18:57:58 | dansmith | and those are somewhat at odds | |
| 18:57:58 | jaypipes | ya | |
| 18:59:23 | mriedem | ok, so again, it's fixing latent bugs which we didn't care about until those latent bugs affected scheduling decisions, which they do now | |
| 18:59:48 | dansmith | they affect ocata scheduler too | |
| 19:00:09 | dansmith | claiming in the scheduler is after the existing decision gets made on the data we're stomping on | |
| 19:00:31 | dansmith | and the stomping just makes placement think there is more room than there is, | |
| 19:00:36 | dansmith | so not claiming doesn't make anything easier | |
| 19:00:44 | dansmith | it makes it less likely to be right, but that's not really useful | |
| 19:01:00 | mriedem | excluding shared storage providers, the computes are eventually consistent aren't they? | |
| 19:01:30 | mriedem | i'll stop asking questions since these are things i've gone in circles on for 2+ weeks now | |
| 19:04:47 | dansmith | I've about got all three patches working together | |
| 19:04:58 | dansmith | which is scary, because I think I was supposed to modify the compute side code to make this work and I haven't done that yet | |
| 19:05:11 | dansmith | (nor do I remember what that was anymore) | |
| 19:05:41 | mriedem | are you fixing the issue in my patch? | |
| 19:06:04 | dansmith | which issue? | |
| 19:06:24 | openstackgerrit | Matt Riedemann proposed openstack/nova master: WIP: Sum allocations in the scheduler when resizing to the same host https://review.openstack.org/490085 | |
| 19:06:25 | openstackgerrit | Matt Riedemann proposed openstack/nova master: WIP: Handle shared storage allocations when resize to same host https://review.openstack.org/490159 | |
| 19:06:30 | mriedem | self.assertFlavorMatcheAllocation | |
| 19:06:37 | mriedem | ^ fixed | |
| 19:07:35 | dansmith | well, I'm just going to push over top of yours | |
| 19:07:38 | dansmith | but yeah I had fixed that | |
| 19:09:57 | dansmith | ooh, nice I'm getting a 500 from placement now | |
| 19:10:40 | cdent | nice work dansmith | |
| 19:11:03 | dansmith | I'll fix this undoubling thing and then push to let you placement peeps look at the 500 | |
| 19:11:23 | cdent | yeah, I can look at that when there are some details | |
| 19:12:58 | mriedem | i'm really confused because https://review.openstack.org/490159 is passing the test i added w/o any code changes to handle it | |
| 19:15:42 | dansmith | mriedem: passing what? it failed gibi's same host tests when I put it on top | |
| 19:16:07 | mriedem | this is the shared storage + resize to same host one | |
| 19:16:18 | mriedem | https://review.openstack.org/#/c/490159/ adds a unit test for that and i expected it to fail | |
| 19:18:17 | mriedem | oh no, i know why it's passing | |
| 19:18:17 | mriedem | heh | |
| 19:18:45 | mriedem | yup | |
| 19:19:15 | mriedem | i'll just squash those changes together | |
| 19:19:27 | dansmith | mriedem: can you hold off? | |
| 19:19:29 | mriedem | cdent: ^ is why i'm looping the allocations | |
| 19:19:35 | mriedem | rather than assuming there is 1 | |
| 19:19:35 | dansmith | I have a bunch of cuts against all three of these patches | |
| 19:19:42 | mriedem | dansmith: like, deep cuts? | |
| 19:19:47 | dansmith | gashes | |
| 19:19:56 | dansmith | with rusty blades | |
| 19:20:10 | mriedem | i meant like https://www.youtube.com/watch?v=KCdKBHdPz30 | |
| 19:20:13 | mriedem | deep cuts | |
| 19:20:18 | dansmith | heh | |
| 19:20:20 | dansmith | nice one | |
| 19:21:12 | mriedem | btw, fagen is forced to tour again https://qz.com/1041397/steely-dans-donald-fagen-is-back-on-tour-the-result-of-nobody-buying-music-albums-anymore/ | |
| 19:21:19 | mriedem | streaming music has broken him | |
| 19:21:50 | mriedem | he should probably talk to a financial advisor about diversifying his portfolio | |
| 19:22:04 | melwitt | lol | |
| 19:22:38 | dfisher | is there a known issue with calling nova.context.get_admin_context() from within a virt driver? http://paste.openstack.org/show/617308/ | |
| 19:23:11 | melwitt | did not expect that article to mention Ronnie James Dio | |
| 19:23:18 | mriedem | hologram dio | |
| 19:23:20 | mriedem | terrible | |
| 19:23:26 | melwitt | I know, like, seriously? | |
| 19:23:49 | mriedem | i was at a sabbath reunion show once and dio scolded the audience and threatened to cut the show and leave if they didn't settle down | |
| 19:24:30 | dansmith | cripes | |
| 19:24:43 | dansmith | jaypipes: how do we get the old_flavor if we're here: https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L460-L460 ? | |
| 19:25:16 | mriedem | dansmith: isn't old_flavor stored on the instance? | |
| 19:25:41 | dansmith | mriedem: well, I would assume not at this point since we go to great lengths to look it up on L443 | |
| 19:26:05 | mriedem | maybe that predates the old_flavor being stored on the instance? | |
| 19:26:08 | dansmith | no, | |
| 19:26:13 | dansmith | we set it to None before we call this here: | |
| 19:26:31 | dansmith | https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L3501-L3521 | |
| 19:26:34 | dansmith | son of a goat | |
| 19:26:41 | mriedem | is it stashed on the migration record? | |
| 19:26:46 | dansmith | no, only the id | |
| 19:26:55 | mriedem | ha | |
| 19:26:55 | mriedem | # NOTE(danms): delete stashed migration information | |
| 19:27:02 | dansmith | wait | |
| 19:27:02 | mriedem | BUT WHY?! | |
| 19:27:20 | dansmith | we pass old_instance_type for instance_type, | |
| 19:27:27 | dansmith | but clearly that can be None sometimes | |
| 19:27:30 | dansmith | but I don't know when | |
| 19:27:36 | dansmith | maybe if we're doing a migrate? | |
| 19:27:43 | mriedem | probably yeah | |
| 19:27:47 | mriedem | no old/new if it's not a resize | |
| 19:27:48 | dansmith | so maybe I can use old_instance_type or instance.flavor | |
| 19:27:54 | mriedem | f yes you can | |
| 19:30:53 | mriedem | dfisher: if you're using a recent devstack, it's running in superconductor mode which means you can't do retries from the compute | |
| 19:31:00 | mriedem | it can't upcall to the api db to get instance group info | |
| 19:31:16 | mriedem | if you're doing some crazy crap in the oracle virt driver, then you're on your own | |
| 19:32:49 | dfisher | i'm really not doing anything crazy … I don't think. | |
| 19:32:59 | dansmith | dfisher: well, except for that one crazy thing | |
| 19:33:18 | dfisher | like, I'm using nova.api.metadata.password.convert_password | |
| 19:33:23 | dfisher | which takes a context | |
| 19:35:07 | dfisher | dansmith: ssssh. don't tell them </loud whisper> | |
| 19:37:38 | mriedem | dfisher: if you're looking at trunk code, build_instances in the conductor is only hit if (1) youre' using cellsv1 which you shouldn't be or (2) if the compute is doing a retry of a build | |
| 19:37:48 | dansmith | okay I think I'm doing the right thing, and now I'm failing because of that 500 from placement | |
| 19:37:56 | dansmith | so I'll push this up and let someone else look at that in parallel | |
| 19:37:59 | mriedem | dfisher: so assuming you're hitting (2), if you're using devstack in the default superconductor mode, the retry isn't going to work | |
| 19:38:16 | dfisher | hmm. ok. | |
| 19:38:20 | mriedem | dfisher: b/c you're hitting the cell conductor which doesn't have access to the scheduler | |
| 19:38:23 | mriedem | or api db for that matter | |