| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-24 | |||
| 13:59:47 | mriedem | https://review.openstack.org/#/c/496995/ | |
| 14:00:04 | mriedem | on build_results.FAILED - so unexpected failures during build, or BuildAbortException | |
| 14:00:09 | mriedem | but, it's meeting time | |
| 14:08:20 | openstackgerrit | Lajos Katona proposed openstack/nova master: WIP: Test server movings with custom resources https://review.openstack.org/497399 | |
| 14:10:32 | cdent | lajoskatona: will take a look at that ^ once the nova meeting is done | |
| 14:15:25 | openstackgerrit | Eric Fried proposed openstack/nova master: Add formatting to scheduling activity diagram https://review.openstack.org/476204 | |
| 14:17:35 | efried | edleafe stephenfin sfinucan ^^ | |
| 14:17:50 | stephenfin | efried: Cheers. Looking | |
| 14:18:20 | efried | stephenfin edleafe Rebased on top of the monkeypatch one and updated for the latest text in the doc. | |
| 14:18:30 | efried | Not too much to see until the docs build completes, I guess :) | |
| 14:20:08 | openstackgerrit | Merged openstack/nova master: How about not logging errors every time we shelve offload? https://review.openstack.org/496930 | |
| 14:20:11 | stephenfin | efried: Yup, that's what I was hoping for. +2d, on the assumption the gate will catch any typos :) | |
| 14:20:56 | dtantsur | hi folks, it's me again :) do you think we should document setting https://github.com/openstack/nova/blob/master/nova/conf/compute.py#L634 to 0 in case of ironic? | |
| 14:21:14 | dtantsur | I don't think the compute instance is a popular cause of build failures for ironic, as opposed to the even world around it | |
| 14:21:58 | efried | dtantsur Are you saying ironic overrides that default today, or are you suggesting that it do so? | |
| 14:22:19 | dtantsur | efried: I'm suggesting we document overriding it for nova-compute processes bound to ironic | |
| 14:22:30 | dtantsur | (maybe in ironic docs, we have a section on configuring nova) | |
| 14:22:38 | dtantsur | I'm just checking if my understanding of it is correct | |
| 14:22:57 | efried | dtantsur Oh, you're suggesting no code changes, just documenting that the user, if using ironic, will want to set it to zero. | |
| 14:23:07 | dtantsur | correct | |
| 14:24:01 | efried | dtantsur So the way I understand that config var, it'll disable the compute host entirely if <N> consecutive builds fail. On the theory that we ought to stop trying to schedule to that host, because it's obviously sick. Why is ironic different in this regard? | |
| 14:24:40 | dtantsur | efried: because the nova-compute hosts merely pipe requests to ironic | |
| 14:25:12 | dtantsur | disabling a compute host will make some part of your ironic fleet unavailable, and that may NOT be the part that caused failures | |
| 14:26:00 | efried | I would think that's always the case ("may NOT be the part that caused failures"). | |
| 14:26:20 | openstackgerrit | Alex Xu proposed openstack/nova master: Remove allocation when booting instance rescheduled or aborted https://review.openstack.org/496995 | |
| 14:26:23 | dtantsur | yeah, usually our failures are caused by network problems or misconfiguration of ironic | |
| 14:26:31 | dtantsur | not by the nova-compute process being broken | |
| 14:26:31 | efried | (btw dtantsur I'm not arguing against what you're suggesting - just playing devil's advocate) | |
| 14:26:42 | dtantsur | yeah, got it :) | |
| 14:27:04 | efried | Well, right; if it's network problems that are persistent enough to cause 10 consecutive failures, you should go fix the problems and then re-enable the compute host. | |
| 14:27:12 | efried | Ditto with misconfiguration, even more so. | |
| 14:27:35 | dtantsur | then why kill nova-compute, but not nova-api? :) | |
| 14:27:50 | dtantsur | both are not the source of the problem, both merely pass the requests on | |
| 14:28:14 | efried | Heh, guess that's a dansmith question. | |
| 14:28:15 | dtantsur | also, not the whole ironic deployment may be broken, but you're disabling a *random* part of it | |
| 14:28:42 | efried | That doesn't seem like an argument specific to ironic. | |
| 14:29:02 | dtantsur | no, because nova-compute for e.g. libvirt is tied to its libvirt instances | |
| 14:29:03 | dansmith | so, | |
| 14:29:11 | dtantsur | while nova-compute for ironic handles some randomly chosen share of ironic nodes | |
| 14:29:16 | dtantsur | which can change with time, btw | |
| 14:29:20 | dansmith | the request for that came with ironic people in the room, talking about such failures, and I'm pretty sure they were on board | |
| 14:29:50 | dansmith | dtantsur: if you fail to build ten things _consecutively_ don't you think that's cause for concern? | |
| 14:29:56 | dtantsur | can we can specific people please then? I'm just seeing it as a source of wrong bug reports coming to me | |
| 14:30:07 | dansmith | like, why not drop that compute out and let the ring rebalance? | |
| 14:30:24 | dtantsur | dansmith: rebalanace annnnndd? fail 10 times more, because the problem was not fixed? | |
| 14:30:31 | dtantsur | and this way until we run out of computes? | |
| 14:30:44 | dansmith | dtantsur: well, it depends on what the problem is of course | |
| 14:31:15 | dtantsur | my point is that with VMs killing a nova-compute process can cut a misbahaving compute node. with ironic it does mostly nothing, until we run out of computes (that's where everything breaks) | |
| 14:31:27 | dansmith | dtantsur: the conversation in the room centered around (a) not letting one compute node be a black hole and (b) making it super obvious that something was wrong | |
| 14:32:00 | dtantsur | this assumes the compute node does *something*. in case of ironic it's a small orchestrator, essentially | |
| 14:32:02 | dansmith | dtantsur: but really, if ironic is failing 100% of the builds, what is the harm? leaving them in isn't going to fix anything | |
| 14:32:41 | dtantsur | it's suprising for users, I think. e.g. they had a networking problem, ironic reported valid failures, they fixed them, and... nothing works | |
| 14:32:44 | dansmith | dtantsur: you can configure two computes differently such that one will behave differently right? like pointing them at two different glance mirrors, and one loses the ability to talk to glance | |
| 14:33:00 | dtantsur | but I cannot do this for ironic though | |
| 14:33:04 | dansmith | dtantsur: you know it's consecutive failures not total failures right? | |
| 14:33:20 | dtantsur | well, when something breaks, it's usually several attempts in a raw | |
| 14:33:41 | dansmith | either way, some documentation on "this may not be the behavior you want for ironic" is fine, but I definitely don't want to just say "for ironic this should be zero" | |
| 14:34:01 | dtantsur | I would be less opinionated here, but I suspect the users are going to see "no valid hosts found" in response, right? | |
| 14:34:15 | dansmith | once all the computes are disabled? yes. | |
| 14:34:17 | cdent | lajoskatona: I’ve responded on that review with some info on what’s going wrong. | |
| 14:34:53 | dansmith | dtantsur: the other thing that helps is that if you're really at 100% fail, eventually you stop wasting time, cpu, and network bandwidth trying to build and prepare things that are going to fail | |
| 14:34:57 | dtantsur | I'm reserving my opinion on this error :) it's too hard to debug, and it can mean anything (especially when RetryFilter is used) | |
| 14:35:18 | dansmith | so getting to zero disabled computes doesn't seem like a bad thing to me if literally 100% of the builds will fail | |
| 14:36:17 | dtantsur | does it include attempts from the RetryFilter? | |
| 14:36:51 | dansmith | I'm not sure what you mean.. are you asking if three retries against the same compute count as three failures? | |
| 14:37:05 | dtantsur | (side note: I became aware of this option after seeing https://review.openstack.org/#/c/496851/ ) | |
| 14:37:11 | dtantsur | dansmith: yep | |
| 14:37:23 | dansmith | dtantsur: sure, it's all the same | |
| 14:38:07 | dtantsur | dansmith: meaning, if people have retry count >= 10, one misconfigured instance can bring down a compute node? | |
| 14:38:29 | dansmith | dtantsur: if they configure retries high and don't adjust this, then sure | |
| 14:39:13 | dtantsur | okay, I think this is at least worth mentioning in our docs, because it may cause surprises (well, it did cause surprises already for tripleo people) | |
| 14:39:46 | openstackgerrit | Ed Leafe proposed openstack/nova master: docs: Document the scheduler workflow https://review.openstack.org/475810 | |
| 14:39:59 | dansmith | dtantsur: what sort of mention do you want? | |
| 14:40:26 | dansmith | note that the text of that config option mentions retries, | |
| 14:40:34 | dansmith | so hopefully someone upping that to a crazy level will see that | |
| 14:41:08 | dtantsur | dansmith: I want to have a mention of it in https://docs.openstack.org/ironic/latest/install/configure-compute.html | |
| 14:41:32 | dtantsur | maybe just something like "Mind the following option, if you don't like <behavior>, change it to 0" | |
| 14:42:15 | dansmith | dtantsur: sure, ironic docs that talk about ideal nova settings when using the two together is certainly a good idea :) | |
| 14:42:43 | dtantsur | yep, that was my initial idea | |
| 14:42:52 | dansmith | I have no problem with that | |
| 14:43:50 | alex_xu | mriedem: would you mind take care of https://review.openstack.org/496995, since the time is late for me | |
| 14:44:34 | dtantsur | cool, thanks all! | |
| 14:45:20 | mriedem | alex_xu: yup, i'll take over, thanks for working on it | |
| 14:45:47 | alex_xu | mriedem: thanks, good luck for rc2 :) | |
| 14:45:49 | mriedem | mnestratov|2: fyi https://bugs.launchpad.net/nova/+bug/1712801 | |
| 14:45:50 | openstack | Launchpad bug 1712801 in OpenStack Compute (nova) "Virtuozzo conainers lack SCSI info in libvirt XML" [Undecided,New] | |
| 14:46:36 | dtantsur | sahid: actually, could you take a quick look at https://docs.openstack.org/ironic/latest/install/configure-compute.html to check if everything is still correct there? there are a few options that I don't remember well | |
| 14:46:40 | dtantsur | oops, wrong ping | |
| 14:46:52 | dtantsur | sorry sahid, I meant dansmith (and please don't ask how I made this typo) | |
| 14:47:08 | sahid | :) | |
| 14:47:55 | mnestratov|2 | mriedem: thanks, this is our guy filed the bug in context of review https://review.openstack.org/#/c/495756/ | |
| 14:49:06 | dansmith | dtantsur: we don't need the hostmanager thing after the resource class transition right? | |
| 14:50:21 | dansmith | I'm also not sure why you tell people to use a crazy host_subset_size without explaining why they may or may not want that | |
| 14:50:36 | dansmith | especially in pike where we shouldn't have any scheduler races | |
| 14:51:12 | dtantsur | I think using a value > 1 is still a good idea, as two instances cannot get on one host | |
| 14:51:28 | dtantsur | probably not so huge, and I agree that we need an explanation | |
| 14:51:52 | dtantsur | and yes, it seems like we should have deprecated ironic_host_manager in pike | |
| 14:52:53 | lajoskatona | cdent: thanks I check it | |
| 14:53:06 | dansmith | dtantsur: with claims in the scheduler and resource classes, you can not have two instances pick the same host | |
| 14:53:18 | efried | edleafe "Return a list of selected host + alternates, along with their allocations to the conductor" <== Is [allocations to the conductor] a thing, or are [a list of selected hosts + alternates, along with their allocations] being returned to the conductor? | |