| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-08 | |||
| 17:21:01 | dansmith | hv stats has never been particularly correct either, mind you | |
| 17:21:36 | melwitt | okay | |
| 17:22:17 | melwitt | I think I could update the test to query placement for how much disk the compute node is reporting | |
| 17:22:43 | melwitt | because that's all I was doing is verifying it's reporting 0 local_gb_used | |
| 17:23:32 | dansmith | so the compute node's allocation will account for it, but we don't heal if we're on pike only, | |
| 17:23:40 | melwitt | it would be interested to see if the test passes with those lines commented out too, because if bfv isn't handled correctly yet, it won't even be able to boot instances | |
| 17:23:50 | melwitt | *interesting | |
| 17:24:03 | dansmith | looking earlier, the place I had added it to the scheduler request didn't seem particularly correct anymore | |
| 17:24:06 | sean-k-mooney | hi everyone quick question is "nova-manage cell_v2 create_cell" updating the db directly or calling the api to create the cell entries? | |
| 17:24:15 | dansmith | sean-k-mooney: directly | |
| 17:24:24 | dansmith | sean-k-mooney: almost all of nova-manage is direct-to-databsae | |
| 17:24:54 | sean-k-mooney | dansmith: cool thanks, that is what i taught. trying debug an issue in kolla and wnated to check where the logs would be | |
| 17:25:26 | dansmith | on the bfv thing, I'm a little concerned about even getting jaypipes' set in before rc at this point, so... | |
| 17:28:44 | melwitt | yeah. I've been trying to mention the bfv thing occasionally throughout this cycle because it's one of the top things people are looking for from resource providers and I haven't been able to tell how close we are to that working | |
| 17:29:30 | melwitt | I'll try stacking the test on top of jay's set locally to see if it passes without the hv stats stuff | |
| 17:30:24 | dansmith | jaypipes' set doesn't affect any of this, AFAIK | |
| 17:30:52 | dansmith | and we're not even accounting for moves properly right now (which is the subject of the set) which is arguably more important | |
| 17:32:05 | dansmith | my zero request patch may actually still work, since we're still calling resources_from_request_spec or whatever it's calls | |
| 17:32:05 | cdent | dansmith: have you got a summary of “what’s wrong today”? It’s hard to keep track. | |
| 17:32:07 | dansmith | *called | |
| 17:32:18 | dansmith | cdent: no | |
| 17:32:20 | openstackgerrit | Doug Hellmann proposed openstack/nova master: use extlinks to build series-specific links https://review.openstack.org/491869 | |
| 17:32:43 | cdent | roger that, I’ll travel with lights | |
| 17:33:39 | melwitt | dansmith: okay. I thought earlier you said we're already accounting for it in the current stuff, I must have misunderstood | |
| 17:34:05 | dansmith | melwitt: I said we are on the compute node, but that the scheduler probably isn't (which I since confirmed) | |
| 17:34:15 | dansmith | and if we're on pike we'll depend on the scheduler having it right | |
| 17:34:26 | melwitt | I see, thanks | |
| 17:34:36 | dansmith | which is what my zero request patch is for.. asking the scheduler for zero root disk | |
| 17:34:38 | openstackgerrit | Spencer Yu proposed openstack/python-novaclient master: Nova client should retry with Retry-After value https://review.openstack.org/447766 | |
| 17:34:50 | melwitt | yeah | |
| 17:35:11 | dansmith | this is somewhat wrapped up in the shared storage thing which we're punting on again, so it really can't be right regardless | |
| 17:36:39 | melwitt | okay. yeah, I was wondering about the shared storage thing too. good to know | |
| 17:38:49 | jaypipes | melwitt: yeah, the bfv accounting issues will most likely remain in Pike unfortunately. | |
| 17:42:27 | dansmith | jaypipes: replied | |
| 17:43:15 | dansmith | jaypipes: I still don't see how we get to the offload thing, and just want that code to be less complicated, but if we're not sure or you've seen it in real life somehow, then we should just leave it in I guess | |
| 17:44:50 | jaypipes | dansmith: this patch is mostly in response to mriedem and alex_xu's comments on https://review.openstack.org/#/c/488510/ ps 27. | |
| 17:45:12 | jaypipes | dansmith: i.e. this from mriedem: | |
| 17:45:12 | jaypipes | Related to #4, there is bug https://bugs.launchpad.net/nova/+bug/1679750 where we don't delete allocations from Placement in nova-api when doing a "local delete", i.e. when the compute host is down, or the instance does not have a host associated (e.g. shelved_offloaded state). | |
| 17:45:13 | openstack | Launchpad bug 1679750 in OpenStack Compute (nova) "Allocations are not cleaned up in placement for instance 'local delete' case" [Medium,In progress] | |
| 17:45:43 | jaypipes | dansmith: but I admit I've probably f'd this series up :( | |
| 17:45:54 | jaypipes | trying to fix these corner cases | |
| 17:46:06 | dansmith | jaypipes: the shelved offloaded thing in that bug is conjecture right? | |
| 17:46:29 | jaypipes | dansmith: yes, but listed by mriedem as something to handle | |
| 17:46:52 | dansmith | jaypipes: yeah, but maybe he didn't dig to see how we'd actually get here :) | |
| 17:47:08 | jaypipes | perhaps. this code is fugly, as you know. | |
| 17:47:35 | mriedem | 3 more times | |
| 17:47:39 | dansmith | jaypipes: that bug as written does't actually say anything about shelved, although I thought mriedem had said something about it | |
| 17:47:43 | openstackgerrit | Merged openstack/nova master: Handle ironicclient failures in Ironic driver https://review.openstack.org/487925 | |
| 17:47:45 | mriedem | i forgot my kid at camp and they called | |
| 17:47:53 | mriedem | so i'll be paying for that in counseling sessions later | |
| 17:48:07 | jaypipes | mriedem: Dad of the Year. | |
| 17:48:11 | mriedem | tbc, i dropped that link in the change as an fyi | |
| 17:48:18 | mriedem | and listed the move operations | |
| 17:48:36 | mriedem | i was not meaning that the bug had to be fixed in that change | |
| 17:48:59 | dansmith | mriedem: well, we're not sure there is a bug | |
| 17:49:02 | dansmith | or, I'm not | |
| 17:49:06 | jaypipes | mriedem: sure. though the changes I put in there were also to highlight code paths that alex_xu had pointed out around evacuate and the local delete problem where the compute host was down. | |
| 17:49:16 | dansmith | if there is, I'm skeptical that offloaded and host==self is a case we can have :) | |
| 17:49:22 | jaypipes | in there == the new bottom patch on the series. | |
| 17:49:35 | dansmith | jaypipes: the local delete path should be deleting the allocation itself right? | |
| 17:49:42 | dansmith | so that there's nothing to clean up here, ideally | |
| 17:49:49 | mriedem | if the instance is offloaded, host is None | |
| 17:50:05 | dansmith | like, we shouldn't actually delete the instance unless we were able to nuke the allocation | |
| 17:50:08 | mriedem | the point of the bug was that yes we should delete allocations in the local delete case, if we care | |
| 17:50:13 | dansmith | right | |
| 17:50:19 | dansmith | so not in RT at all, IMHO | |
| 17:50:30 | jaypipes | dansmith: the local delete path never gets to the compute host, thus the need for the allocation to be deleted when the compute host starts back up and sees allocations for instances that no longer exist. | |
| 17:50:40 | mriedem | melwitt wrote a test that shows that even though we don't delete the allocations in the api in local delete cases, when the compute host comes back up and the instance is gone, the allocation is deleted by the compute | |
| 17:50:44 | dansmith | jaypipes: no dude, delete it from the api :) | |
| 17:50:47 | mriedem | now ^ might be impacted by whatever you guys are doing | |
| 17:50:58 | dansmith | jaypipes: and don't mark the instance as deleted until you succeeded, or there isn't an allocation | |
| 17:51:08 | mriedem | this https://review.openstack.org/#/c/470578/ | |
| 17:51:09 | dansmith | jaypipes: then there's less complexity for the compute node to handle | |
| 17:51:33 | jaypipes | dansmith: oh, you mean delete the allocation from placement during the nova-api's "local delete" operation? | |
| 17:52:02 | jaypipes | dansmith: we'd still need to deal with ocata apis that didn't do that though ;) yay. | |
| 17:52:03 | dansmith | jaypipes: yeah, we delete anything we can from there without the compute node, | |
| 17:52:04 | dansmith | which is kinda the point of it | |
| 17:52:10 | dansmith | since we can totes nuke the allocation we're good | |
| 17:52:16 | dansmith | jaypipes: why? | |
| 17:52:28 | dansmith | jaypipes: if we nuke it from api, we have to be pike, an ocata compute isn't going to care, right? | |
| 17:52:36 | jaypipes | dansmith: are you saying we *currently* delete the allocation from the API? | |
| 17:52:51 | dansmith | no I'm saying we should do that, and that's what the bug is about | |
| 17:53:04 | dansmith | mriedem: right? | |
| 17:54:00 | jaypipes | dansmith: if an ocata api local-deleted an instance, then is upgraded to pike, the compute host is upgraded to pike as well, wouldn't there be an allocation left over for the compute host that would need deleting? | |
| 17:54:36 | mriedem | melwitt: question in https://review.openstack.org/#/c/470578/ | |
| 17:54:52 | dansmith | jaypipes: sure, but you're doing that right? | |
| 17:54:58 | melwitt | mriedem: I'll add console proxy stuff to the cells docs. I was thinking to mention that "in the future" we're planning to change the location, so people have a heads up | |
| 17:55:10 | mriedem | dansmith: correct | |
| 17:55:30 | melwitt | mriedem: also thanks for rebasing those backport series. didn't even get a chance to ask you yet :) | |
| 17:55:30 | mriedem | during local delete in the API, we delete shit from external services like cinder/neutron because the compute is down or the instance doesn't have a host (it's offloaded) | |
| 17:55:37 | mriedem | placement would be included in "external shit" | |
| 17:55:42 | jaypipes | dansmith: heh, yes, I am. sorry, I thought you were saying there'd be no need for that code. | |
| 17:56:22 | dansmith | jaypipes: no, I don't think there's a need for the shelve offload part I commented on, but for local delete we should hope to never hit this code if we deleted from the api as expected | |
| 17:56:30 | jaypipes | ahhh, sorry. | |
| 17:57:08 | dansmith | jaypipes: if there is an allocation for us that refers to an instance that is marked as deleted, then we should delete the allocation | |
| 17:57:19 | mriedem | when we shelve offload an instance, the allocations for that compute node should be cleaned up by the RT | |
| 17:57:24 | dansmith | jaypipes: note my comment about notfound for later though | |
| 17:57:26 | melwitt | mriedem: to your question, yeah it seems like it would be racy. not sure what else to do though. | |
| 17:57:43 | dansmith | mriedem: we clean the allocations before we offload it, so I don't think we need to handle cleanup there | |
| 17:57:58 | mriedem | dansmith: via RT yeah? | |