Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-24
13:50:46 mriedem that gets all allocations for the node
13:51:09 mriedem and if we get InstanceNotFound looking up the instance, it deletes the allocations for that instance
13:55:06 mriedem left some more comments
13:56:05 mriedem probably need to get dansmith's opinion when he's around. i think we could go with this patch even though the periodic would cleanup the allocations for the node once the instance is deleted, and if we removed the allocations from the api during local delete that would also take care of it - but we don't have that fix yet
13:57:08 dansmith about whether or not we should delete allocations on local delete/
13:59:43 mriedem no,
13:59:47 mriedem https://review.openstack.org/#/c/496995/
14:00:04 mriedem on build_results.FAILED - so unexpected failures during build, or BuildAbortException
14:00:09 mriedem but, it's meeting time
14:08:20 openstackgerrit Lajos Katona proposed openstack/nova master: WIP: Test server movings with custom resources https://review.openstack.org/497399
14:10:32 cdent lajoskatona: will take a look at that ^ once the nova meeting is done
14:15:25 openstackgerrit Eric Fried proposed openstack/nova master: Add formatting to scheduling activity diagram https://review.openstack.org/476204
14:17:35 efried edleafe stephenfin sfinucan ^^
14:17:50 stephenfin efried: Cheers. Looking
14:18:20 efried stephenfin edleafe Rebased on top of the monkeypatch one and updated for the latest text in the doc.
14:18:30 efried Not too much to see until the docs build completes, I guess :)
14:20:08 openstackgerrit Merged openstack/nova master: How about not logging errors every time we shelve offload? https://review.openstack.org/496930
14:20:11 stephenfin efried: Yup, that's what I was hoping for. +2d, on the assumption the gate will catch any typos :)
14:20:56 dtantsur hi folks, it's me again :) do you think we should document setting https://github.com/openstack/nova/blob/master/nova/conf/compute.py#L634 to 0 in case of ironic?
14:21:14 dtantsur I don't think the compute instance is a popular cause of build failures for ironic, as opposed to the even world around it
14:21:58 efried dtantsur Are you saying ironic overrides that default today, or are you suggesting that it do so?
14:22:19 dtantsur efried: I'm suggesting we document overriding it for nova-compute processes bound to ironic
14:22:30 dtantsur (maybe in ironic docs, we have a section on configuring nova)
14:22:38 dtantsur I'm just checking if my understanding of it is correct
14:22:57 efried dtantsur Oh, you're suggesting no code changes, just documenting that the user, if using ironic, will want to set it to zero.
14:23:07 dtantsur correct
14:24:01 efried dtantsur So the way I understand that config var, it'll disable the compute host entirely if <N> consecutive builds fail. On the theory that we ought to stop trying to schedule to that host, because it's obviously sick. Why is ironic different in this regard?
14:24:40 dtantsur efried: because the nova-compute hosts merely pipe requests to ironic
14:25:12 dtantsur disabling a compute host will make some part of your ironic fleet unavailable, and that may NOT be the part that caused failures
14:26:00 efried I would think that's always the case ("may NOT be the part that caused failures").
14:26:20 openstackgerrit Alex Xu proposed openstack/nova master: Remove allocation when booting instance rescheduled or aborted https://review.openstack.org/496995
14:26:23 dtantsur yeah, usually our failures are caused by network problems or misconfiguration of ironic
14:26:31 dtantsur not by the nova-compute process being broken
14:26:31 efried (btw dtantsur I'm not arguing against what you're suggesting - just playing devil's advocate)
14:26:42 dtantsur yeah, got it :)
14:27:04 efried Well, right; if it's network problems that are persistent enough to cause 10 consecutive failures, you should go fix the problems and then re-enable the compute host.
14:27:12 efried Ditto with misconfiguration, even more so.
14:27:35 dtantsur then why kill nova-compute, but not nova-api? :)
14:27:50 dtantsur both are not the source of the problem, both merely pass the requests on
14:28:14 efried Heh, guess that's a dansmith question.
14:28:15 dtantsur also, not the whole ironic deployment may be broken, but you're disabling a *random* part of it
14:28:42 efried That doesn't seem like an argument specific to ironic.
14:29:02 dtantsur no, because nova-compute for e.g. libvirt is tied to its libvirt instances
14:29:03 dansmith so,
14:29:11 dtantsur while nova-compute for ironic handles some randomly chosen share of ironic nodes
14:29:16 dtantsur which can change with time, btw
14:29:20 dansmith the request for that came with ironic people in the room, talking about such failures, and I'm pretty sure they were on board
14:29:50 dansmith dtantsur: if you fail to build ten things _consecutively_ don't you think that's cause for concern?
14:29:56 dtantsur can we can specific people please then? I'm just seeing it as a source of wrong bug reports coming to me
14:30:07 dansmith like, why not drop that compute out and let the ring rebalance?
14:30:24 dtantsur dansmith: rebalanace annnnndd? fail 10 times more, because the problem was not fixed?
14:30:31 dtantsur and this way until we run out of computes?
14:30:44 dansmith dtantsur: well, it depends on what the problem is of course
14:31:15 dtantsur my point is that with VMs killing a nova-compute process can cut a misbahaving compute node. with ironic it does mostly nothing, until we run out of computes (that's where everything breaks)
14:31:27 dansmith dtantsur: the conversation in the room centered around (a) not letting one compute node be a black hole and (b) making it super obvious that something was wrong
14:32:00 dtantsur this assumes the compute node does *something*. in case of ironic it's a small orchestrator, essentially
14:32:02 dansmith dtantsur: but really, if ironic is failing 100% of the builds, what is the harm? leaving them in isn't going to fix anything
14:32:41 dtantsur it's suprising for users, I think. e.g. they had a networking problem, ironic reported valid failures, they fixed them, and... nothing works
14:32:44 dansmith dtantsur: you can configure two computes differently such that one will behave differently right? like pointing them at two different glance mirrors, and one loses the ability to talk to glance
14:33:00 dtantsur but I cannot do this for ironic though
14:33:04 dansmith dtantsur: you know it's consecutive failures not total failures right?
14:33:20 dtantsur well, when something breaks, it's usually several attempts in a raw
14:33:41 dansmith either way, some documentation on "this may not be the behavior you want for ironic" is fine, but I definitely don't want to just say "for ironic this should be zero"
14:34:01 dtantsur I would be less opinionated here, but I suspect the users are going to see "no valid hosts found" in response, right?
14:34:15 dansmith once all the computes are disabled? yes.
14:34:17 cdent lajoskatona: I’ve responded on that review with some info on what’s going wrong.
14:34:53 dansmith dtantsur: the other thing that helps is that if you're really at 100% fail, eventually you stop wasting time, cpu, and network bandwidth trying to build and prepare things that are going to fail
14:34:57 dtantsur I'm reserving my opinion on this error :) it's too hard to debug, and it can mean anything (especially when RetryFilter is used)
14:35:18 dansmith so getting to zero disabled computes doesn't seem like a bad thing to me if literally 100% of the builds will fail
14:36:17 dtantsur does it include attempts from the RetryFilter?
14:36:51 dansmith I'm not sure what you mean.. are you asking if three retries against the same compute count as three failures?
14:37:05 dtantsur (side note: I became aware of this option after seeing https://review.openstack.org/#/c/496851/ )
14:37:11 dtantsur dansmith: yep
14:37:23 dansmith dtantsur: sure, it's all the same
14:38:07 dtantsur dansmith: meaning, if people have retry count >= 10, one misconfigured instance can bring down a compute node?
14:38:29 dansmith dtantsur: if they configure retries high and don't adjust this, then sure
14:39:13 dtantsur okay, I think this is at least worth mentioning in our docs, because it may cause surprises (well, it did cause surprises already for tripleo people)
14:39:46 openstackgerrit Ed Leafe proposed openstack/nova master: docs: Document the scheduler workflow https://review.openstack.org/475810
14:39:59 dansmith dtantsur: what sort of mention do you want?
14:40:26 dansmith note that the text of that config option mentions retries,
14:40:34 dansmith so hopefully someone upping that to a crazy level will see that
14:41:08 dtantsur dansmith: I want to have a mention of it in https://docs.openstack.org/ironic/latest/install/configure-compute.html
14:41:32 dtantsur maybe just something like "Mind the following option, if you don't like <behavior>, change it to 0"
14:42:15 dansmith dtantsur: sure, ironic docs that talk about ideal nova settings when using the two together is certainly a good idea :)
14:42:43 dtantsur yep, that was my initial idea
14:42:52 dansmith I have no problem with that
14:43:50 alex_xu mriedem: would you mind take care of https://review.openstack.org/496995, since the time is late for me
14:44:34 dtantsur cool, thanks all!
14:45:20 mriedem alex_xu: yup, i'll take over, thanks for working on it
14:45:47 alex_xu mriedem: thanks, good luck for rc2 :)
14:45:49 mriedem mnestratov|2: fyi https://bugs.launchpad.net/nova/+bug/1712801
14:45:50 openstack Launchpad bug 1712801 in OpenStack Compute (nova) "Virtuozzo conainers lack SCSI info in libvirt XML" [Undecided,New]
14:46:36 dtantsur sahid: actually, could you take a quick look at https://docs.openstack.org/ironic/latest/install/configure-compute.html to check if everything is still correct there? there are a few options that I don't remember well
14:46:40 dtantsur oops, wrong ping
14:46:52 dtantsur sorry sahid, I meant dansmith (and please don't ask how I made this typo)
14:47:08 sahid :)
14:47:55 mnestratov|2 mriedem: thanks, this is our guy filed the bug in context of review https://review.openstack.org/#/c/495756/
14:49:06 dansmith dtantsur: we don't need the hostmanager thing after the resource class transition right?
14:50:21 dansmith I'm also not sure why you tell people to use a crazy host_subset_size without explaining why they may or may not want that
14:50:36 dansmith especially in pike where we shouldn't have any scheduler races

Earlier   Later