Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-24
13:44:03 lajoskatona cdent: shall I upload as WIP, and somebody should check it what am I missing from the workflow?
13:44:03 cdent can you push up your code as a wip so it is easier to look at?
13:44:06 cdent :)
13:44:11 cdent yes!
13:44:16 lajoskatona cdent, ok, I upload :-)
13:44:23 alex_xu mriedem: cdent in the original RT behaviour, the instance claim will release the resource, we also nill out instance's host and name, that also means the resource will be released by RT
13:45:32 mriedem i'm replying
13:47:51 mriedem guh
13:47:54 mriedem ok replied
13:48:07 mriedem alex_xu: so you've got a point, as usual :)
13:48:28 mriedem if the instance.host is None, then the API isn't going to cast to the compute to complete the deletion which would remove the allocation
13:48:39 mriedem and we don't currently remove the allocations in the API during a 'local' delete
13:48:45 mriedem so we're kind of stuck
13:49:02 cdent is there a reason we don’t do it from the API?
13:49:09 cdent s/don’t/don’t want to/
13:49:09 mriedem hasn't been fixed yet
13:49:17 openstack Launchpad bug 1679750 in OpenStack Compute (nova) "Allocations are not cleaned up in placement for instance 'local delete' case" [Medium,Confirmed]
13:49:17 mriedem no, it's just an open bug https://bugs.launchpad.net/nova/+bug/1679750
13:49:37 mriedem cdent: it was less of a problem at the time it was reported because the periodic task would heal and remove the allocations
13:50:02 mriedem well, actually let me see about that
13:50:41 mriedem yeah, actually the periodic task should still clean these up, in _remove_deleted_instances_allocations
13:50:46 mriedem that gets all allocations for the node
13:51:09 mriedem and if we get InstanceNotFound looking up the instance, it deletes the allocations for that instance
13:55:06 mriedem left some more comments
13:56:05 mriedem probably need to get dansmith's opinion when he's around. i think we could go with this patch even though the periodic would cleanup the allocations for the node once the instance is deleted, and if we removed the allocations from the api during local delete that would also take care of it - but we don't have that fix yet
13:57:08 dansmith about whether or not we should delete allocations on local delete/
13:59:43 mriedem no,
13:59:47 mriedem https://review.openstack.org/#/c/496995/
14:00:04 mriedem on build_results.FAILED - so unexpected failures during build, or BuildAbortException
14:00:09 mriedem but, it's meeting time
14:08:20 openstackgerrit Lajos Katona proposed openstack/nova master: WIP: Test server movings with custom resources https://review.openstack.org/497399
14:10:32 cdent lajoskatona: will take a look at that ^ once the nova meeting is done
14:15:25 openstackgerrit Eric Fried proposed openstack/nova master: Add formatting to scheduling activity diagram https://review.openstack.org/476204
14:17:35 efried edleafe stephenfin sfinucan ^^
14:17:50 stephenfin efried: Cheers. Looking
14:18:20 efried stephenfin edleafe Rebased on top of the monkeypatch one and updated for the latest text in the doc.
14:18:30 efried Not too much to see until the docs build completes, I guess :)
14:20:08 openstackgerrit Merged openstack/nova master: How about not logging errors every time we shelve offload? https://review.openstack.org/496930
14:20:11 stephenfin efried: Yup, that's what I was hoping for. +2d, on the assumption the gate will catch any typos :)
14:20:56 dtantsur hi folks, it's me again :) do you think we should document setting https://github.com/openstack/nova/blob/master/nova/conf/compute.py#L634 to 0 in case of ironic?
14:21:14 dtantsur I don't think the compute instance is a popular cause of build failures for ironic, as opposed to the even world around it
14:21:58 efried dtantsur Are you saying ironic overrides that default today, or are you suggesting that it do so?
14:22:19 dtantsur efried: I'm suggesting we document overriding it for nova-compute processes bound to ironic
14:22:30 dtantsur (maybe in ironic docs, we have a section on configuring nova)
14:22:38 dtantsur I'm just checking if my understanding of it is correct
14:22:57 efried dtantsur Oh, you're suggesting no code changes, just documenting that the user, if using ironic, will want to set it to zero.
14:23:07 dtantsur correct
14:24:01 efried dtantsur So the way I understand that config var, it'll disable the compute host entirely if <N> consecutive builds fail. On the theory that we ought to stop trying to schedule to that host, because it's obviously sick. Why is ironic different in this regard?
14:24:40 dtantsur efried: because the nova-compute hosts merely pipe requests to ironic
14:25:12 dtantsur disabling a compute host will make some part of your ironic fleet unavailable, and that may NOT be the part that caused failures
14:26:00 efried I would think that's always the case ("may NOT be the part that caused failures").
14:26:20 openstackgerrit Alex Xu proposed openstack/nova master: Remove allocation when booting instance rescheduled or aborted https://review.openstack.org/496995
14:26:23 dtantsur yeah, usually our failures are caused by network problems or misconfiguration of ironic
14:26:31 efried (btw dtantsur I'm not arguing against what you're suggesting - just playing devil's advocate)
14:26:31 dtantsur not by the nova-compute process being broken
14:26:42 dtantsur yeah, got it :)
14:27:04 efried Well, right; if it's network problems that are persistent enough to cause 10 consecutive failures, you should go fix the problems and then re-enable the compute host.
14:27:12 efried Ditto with misconfiguration, even more so.
14:27:35 dtantsur then why kill nova-compute, but not nova-api? :)
14:27:50 dtantsur both are not the source of the problem, both merely pass the requests on
14:28:14 efried Heh, guess that's a dansmith question.
14:28:15 dtantsur also, not the whole ironic deployment may be broken, but you're disabling a *random* part of it
14:28:42 efried That doesn't seem like an argument specific to ironic.
14:29:02 dtantsur no, because nova-compute for e.g. libvirt is tied to its libvirt instances
14:29:03 dansmith so,
14:29:11 dtantsur while nova-compute for ironic handles some randomly chosen share of ironic nodes
14:29:16 dtantsur which can change with time, btw
14:29:20 dansmith the request for that came with ironic people in the room, talking about such failures, and I'm pretty sure they were on board
14:29:50 dansmith dtantsur: if you fail to build ten things _consecutively_ don't you think that's cause for concern?
14:29:56 dtantsur can we can specific people please then? I'm just seeing it as a source of wrong bug reports coming to me
14:30:07 dansmith like, why not drop that compute out and let the ring rebalance?
14:30:24 dtantsur dansmith: rebalanace annnnndd? fail 10 times more, because the problem was not fixed?
14:30:31 dtantsur and this way until we run out of computes?
14:30:44 dansmith dtantsur: well, it depends on what the problem is of course
14:31:15 dtantsur my point is that with VMs killing a nova-compute process can cut a misbahaving compute node. with ironic it does mostly nothing, until we run out of computes (that's where everything breaks)
14:31:27 dansmith dtantsur: the conversation in the room centered around (a) not letting one compute node be a black hole and (b) making it super obvious that something was wrong
14:32:00 dtantsur this assumes the compute node does *something*. in case of ironic it's a small orchestrator, essentially
14:32:02 dansmith dtantsur: but really, if ironic is failing 100% of the builds, what is the harm? leaving them in isn't going to fix anything
14:32:41 dtantsur it's suprising for users, I think. e.g. they had a networking problem, ironic reported valid failures, they fixed them, and... nothing works
14:32:44 dansmith dtantsur: you can configure two computes differently such that one will behave differently right? like pointing them at two different glance mirrors, and one loses the ability to talk to glance
14:33:00 dtantsur but I cannot do this for ironic though
14:33:04 dansmith dtantsur: you know it's consecutive failures not total failures right?
14:33:20 dtantsur well, when something breaks, it's usually several attempts in a raw
14:33:41 dansmith either way, some documentation on "this may not be the behavior you want for ironic" is fine, but I definitely don't want to just say "for ironic this should be zero"
14:34:01 dtantsur I would be less opinionated here, but I suspect the users are going to see "no valid hosts found" in response, right?
14:34:15 dansmith once all the computes are disabled? yes.
14:34:17 cdent lajoskatona: I’ve responded on that review with some info on what’s going wrong.
14:34:53 dansmith dtantsur: the other thing that helps is that if you're really at 100% fail, eventually you stop wasting time, cpu, and network bandwidth trying to build and prepare things that are going to fail
14:34:57 dtantsur I'm reserving my opinion on this error :) it's too hard to debug, and it can mean anything (especially when RetryFilter is used)
14:35:18 dansmith so getting to zero disabled computes doesn't seem like a bad thing to me if literally 100% of the builds will fail
14:36:17 dtantsur does it include attempts from the RetryFilter?
14:36:51 dansmith I'm not sure what you mean.. are you asking if three retries against the same compute count as three failures?
14:37:05 dtantsur (side note: I became aware of this option after seeing https://review.openstack.org/#/c/496851/ )
14:37:11 dtantsur dansmith: yep
14:37:23 dansmith dtantsur: sure, it's all the same
14:38:07 dtantsur dansmith: meaning, if people have retry count >= 10, one misconfigured instance can bring down a compute node?
14:38:29 dansmith dtantsur: if they configure retries high and don't adjust this, then sure
14:39:13 dtantsur okay, I think this is at least worth mentioning in our docs, because it may cause surprises (well, it did cause surprises already for tripleo people)
14:39:46 openstackgerrit Ed Leafe proposed openstack/nova master: docs: Document the scheduler workflow https://review.openstack.org/475810
14:39:59 dansmith dtantsur: what sort of mention do you want?

Earlier   Later