| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-24 | |||
| 13:48:28 | mriedem | if the instance.host is None, then the API isn't going to cast to the compute to complete the deletion which would remove the allocation | |
| 13:48:39 | mriedem | and we don't currently remove the allocations in the API during a 'local' delete | |
| 13:48:45 | mriedem | so we're kind of stuck | |
| 13:49:02 | cdent | is there a reason we don’t do it from the API? | |
| 13:49:09 | mriedem | hasn't been fixed yet | |
| 13:49:09 | cdent | s/don’t/don’t want to/ | |
| 13:49:17 | mriedem | no, it's just an open bug https://bugs.launchpad.net/nova/+bug/1679750 | |
| 13:49:17 | openstack | Launchpad bug 1679750 in OpenStack Compute (nova) "Allocations are not cleaned up in placement for instance 'local delete' case" [Medium,Confirmed] | |
| 13:49:37 | mriedem | cdent: it was less of a problem at the time it was reported because the periodic task would heal and remove the allocations | |
| 13:50:02 | mriedem | well, actually let me see about that | |
| 13:50:41 | mriedem | yeah, actually the periodic task should still clean these up, in _remove_deleted_instances_allocations | |
| 13:50:46 | mriedem | that gets all allocations for the node | |
| 13:51:09 | mriedem | and if we get InstanceNotFound looking up the instance, it deletes the allocations for that instance | |
| 13:55:06 | mriedem | left some more comments | |
| 13:56:05 | mriedem | probably need to get dansmith's opinion when he's around. i think we could go with this patch even though the periodic would cleanup the allocations for the node once the instance is deleted, and if we removed the allocations from the api during local delete that would also take care of it - but we don't have that fix yet | |
| 13:57:08 | dansmith | about whether or not we should delete allocations on local delete/ | |
| 13:59:43 | mriedem | no, | |
| 13:59:47 | mriedem | https://review.openstack.org/#/c/496995/ | |
| 14:00:04 | mriedem | on build_results.FAILED - so unexpected failures during build, or BuildAbortException | |
| 14:00:09 | mriedem | but, it's meeting time | |
| 14:08:20 | openstackgerrit | Lajos Katona proposed openstack/nova master: WIP: Test server movings with custom resources https://review.openstack.org/497399 | |
| 14:10:32 | cdent | lajoskatona: will take a look at that ^ once the nova meeting is done | |
| 14:15:25 | openstackgerrit | Eric Fried proposed openstack/nova master: Add formatting to scheduling activity diagram https://review.openstack.org/476204 | |
| 14:17:35 | efried | edleafe stephenfin sfinucan ^^ | |
| 14:17:50 | stephenfin | efried: Cheers. Looking | |
| 14:18:20 | efried | stephenfin edleafe Rebased on top of the monkeypatch one and updated for the latest text in the doc. | |
| 14:18:30 | efried | Not too much to see until the docs build completes, I guess :) | |
| 14:20:08 | openstackgerrit | Merged openstack/nova master: How about not logging errors every time we shelve offload? https://review.openstack.org/496930 | |
| 14:20:11 | stephenfin | efried: Yup, that's what I was hoping for. +2d, on the assumption the gate will catch any typos :) | |
| 14:20:56 | dtantsur | hi folks, it's me again :) do you think we should document setting https://github.com/openstack/nova/blob/master/nova/conf/compute.py#L634 to 0 in case of ironic? | |
| 14:21:14 | dtantsur | I don't think the compute instance is a popular cause of build failures for ironic, as opposed to the even world around it | |
| 14:21:58 | efried | dtantsur Are you saying ironic overrides that default today, or are you suggesting that it do so? | |
| 14:22:19 | dtantsur | efried: I'm suggesting we document overriding it for nova-compute processes bound to ironic | |
| 14:22:30 | dtantsur | (maybe in ironic docs, we have a section on configuring nova) | |
| 14:22:38 | dtantsur | I'm just checking if my understanding of it is correct | |
| 14:22:57 | efried | dtantsur Oh, you're suggesting no code changes, just documenting that the user, if using ironic, will want to set it to zero. | |
| 14:23:07 | dtantsur | correct | |
| 14:24:01 | efried | dtantsur So the way I understand that config var, it'll disable the compute host entirely if <N> consecutive builds fail. On the theory that we ought to stop trying to schedule to that host, because it's obviously sick. Why is ironic different in this regard? | |
| 14:24:40 | dtantsur | efried: because the nova-compute hosts merely pipe requests to ironic | |
| 14:25:12 | dtantsur | disabling a compute host will make some part of your ironic fleet unavailable, and that may NOT be the part that caused failures | |
| 14:26:00 | efried | I would think that's always the case ("may NOT be the part that caused failures"). | |
| 14:26:20 | openstackgerrit | Alex Xu proposed openstack/nova master: Remove allocation when booting instance rescheduled or aborted https://review.openstack.org/496995 | |
| 14:26:23 | dtantsur | yeah, usually our failures are caused by network problems or misconfiguration of ironic | |
| 14:26:31 | dtantsur | not by the nova-compute process being broken | |
| 14:26:31 | efried | (btw dtantsur I'm not arguing against what you're suggesting - just playing devil's advocate) | |
| 14:26:42 | dtantsur | yeah, got it :) | |
| 14:27:04 | efried | Well, right; if it's network problems that are persistent enough to cause 10 consecutive failures, you should go fix the problems and then re-enable the compute host. | |
| 14:27:12 | efried | Ditto with misconfiguration, even more so. | |
| 14:27:35 | dtantsur | then why kill nova-compute, but not nova-api? :) | |
| 14:27:50 | dtantsur | both are not the source of the problem, both merely pass the requests on | |
| 14:28:14 | efried | Heh, guess that's a dansmith question. | |
| 14:28:15 | dtantsur | also, not the whole ironic deployment may be broken, but you're disabling a *random* part of it | |
| 14:28:42 | efried | That doesn't seem like an argument specific to ironic. | |
| 14:29:02 | dtantsur | no, because nova-compute for e.g. libvirt is tied to its libvirt instances | |
| 14:29:03 | dansmith | so, | |
| 14:29:11 | dtantsur | while nova-compute for ironic handles some randomly chosen share of ironic nodes | |
| 14:29:16 | dtantsur | which can change with time, btw | |
| 14:29:20 | dansmith | the request for that came with ironic people in the room, talking about such failures, and I'm pretty sure they were on board | |
| 14:29:50 | dansmith | dtantsur: if you fail to build ten things _consecutively_ don't you think that's cause for concern? | |
| 14:29:56 | dtantsur | can we can specific people please then? I'm just seeing it as a source of wrong bug reports coming to me | |
| 14:30:07 | dansmith | like, why not drop that compute out and let the ring rebalance? | |
| 14:30:24 | dtantsur | dansmith: rebalanace annnnndd? fail 10 times more, because the problem was not fixed? | |
| 14:30:31 | dtantsur | and this way until we run out of computes? | |
| 14:30:44 | dansmith | dtantsur: well, it depends on what the problem is of course | |
| 14:31:15 | dtantsur | my point is that with VMs killing a nova-compute process can cut a misbahaving compute node. with ironic it does mostly nothing, until we run out of computes (that's where everything breaks) | |
| 14:31:27 | dansmith | dtantsur: the conversation in the room centered around (a) not letting one compute node be a black hole and (b) making it super obvious that something was wrong | |
| 14:32:00 | dtantsur | this assumes the compute node does *something*. in case of ironic it's a small orchestrator, essentially | |
| 14:32:02 | dansmith | dtantsur: but really, if ironic is failing 100% of the builds, what is the harm? leaving them in isn't going to fix anything | |
| 14:32:41 | dtantsur | it's suprising for users, I think. e.g. they had a networking problem, ironic reported valid failures, they fixed them, and... nothing works | |
| 14:32:44 | dansmith | dtantsur: you can configure two computes differently such that one will behave differently right? like pointing them at two different glance mirrors, and one loses the ability to talk to glance | |
| 14:33:00 | dtantsur | but I cannot do this for ironic though | |
| 14:33:04 | dansmith | dtantsur: you know it's consecutive failures not total failures right? | |
| 14:33:20 | dtantsur | well, when something breaks, it's usually several attempts in a raw | |
| 14:33:41 | dansmith | either way, some documentation on "this may not be the behavior you want for ironic" is fine, but I definitely don't want to just say "for ironic this should be zero" | |
| 14:34:01 | dtantsur | I would be less opinionated here, but I suspect the users are going to see "no valid hosts found" in response, right? | |
| 14:34:15 | dansmith | once all the computes are disabled? yes. | |
| 14:34:17 | cdent | lajoskatona: I’ve responded on that review with some info on what’s going wrong. | |
| 14:34:53 | dansmith | dtantsur: the other thing that helps is that if you're really at 100% fail, eventually you stop wasting time, cpu, and network bandwidth trying to build and prepare things that are going to fail | |
| 14:34:57 | dtantsur | I'm reserving my opinion on this error :) it's too hard to debug, and it can mean anything (especially when RetryFilter is used) | |
| 14:35:18 | dansmith | so getting to zero disabled computes doesn't seem like a bad thing to me if literally 100% of the builds will fail | |
| 14:36:17 | dtantsur | does it include attempts from the RetryFilter? | |
| 14:36:51 | dansmith | I'm not sure what you mean.. are you asking if three retries against the same compute count as three failures? | |
| 14:37:05 | dtantsur | (side note: I became aware of this option after seeing https://review.openstack.org/#/c/496851/ ) | |
| 14:37:11 | dtantsur | dansmith: yep | |
| 14:37:23 | dansmith | dtantsur: sure, it's all the same | |
| 14:38:07 | dtantsur | dansmith: meaning, if people have retry count >= 10, one misconfigured instance can bring down a compute node? | |
| 14:38:29 | dansmith | dtantsur: if they configure retries high and don't adjust this, then sure | |
| 14:39:13 | dtantsur | okay, I think this is at least worth mentioning in our docs, because it may cause surprises (well, it did cause surprises already for tripleo people) | |
| 14:39:46 | openstackgerrit | Ed Leafe proposed openstack/nova master: docs: Document the scheduler workflow https://review.openstack.org/475810 | |
| 14:39:59 | dansmith | dtantsur: what sort of mention do you want? | |
| 14:40:26 | dansmith | note that the text of that config option mentions retries, | |
| 14:40:34 | dansmith | so hopefully someone upping that to a crazy level will see that | |
| 14:41:08 | dtantsur | dansmith: I want to have a mention of it in https://docs.openstack.org/ironic/latest/install/configure-compute.html | |
| 14:41:32 | dtantsur | maybe just something like "Mind the following option, if you don't like <behavior>, change it to 0" | |
| 14:42:15 | dansmith | dtantsur: sure, ironic docs that talk about ideal nova settings when using the two together is certainly a good idea :) | |
| 14:42:43 | dtantsur | yep, that was my initial idea | |
| 14:42:52 | dansmith | I have no problem with that | |
| 14:43:50 | alex_xu | mriedem: would you mind take care of https://review.openstack.org/496995, since the time is late for me | |
| 14:44:34 | dtantsur | cool, thanks all! | |
| 14:45:20 | mriedem | alex_xu: yup, i'll take over, thanks for working on it | |