| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-24 | |||
| 14:30:24 | dtantsur | dansmith: rebalanace annnnndd? fail 10 times more, because the problem was not fixed? | |
| 14:30:31 | dtantsur | and this way until we run out of computes? | |
| 14:30:44 | dansmith | dtantsur: well, it depends on what the problem is of course | |
| 14:31:15 | dtantsur | my point is that with VMs killing a nova-compute process can cut a misbahaving compute node. with ironic it does mostly nothing, until we run out of computes (that's where everything breaks) | |
| 14:31:27 | dansmith | dtantsur: the conversation in the room centered around (a) not letting one compute node be a black hole and (b) making it super obvious that something was wrong | |
| 14:32:00 | dtantsur | this assumes the compute node does *something*. in case of ironic it's a small orchestrator, essentially | |
| 14:32:02 | dansmith | dtantsur: but really, if ironic is failing 100% of the builds, what is the harm? leaving them in isn't going to fix anything | |
| 14:32:41 | dtantsur | it's suprising for users, I think. e.g. they had a networking problem, ironic reported valid failures, they fixed them, and... nothing works | |
| 14:32:44 | dansmith | dtantsur: you can configure two computes differently such that one will behave differently right? like pointing them at two different glance mirrors, and one loses the ability to talk to glance | |
| 14:33:00 | dtantsur | but I cannot do this for ironic though | |
| 14:33:04 | dansmith | dtantsur: you know it's consecutive failures not total failures right? | |
| 14:33:20 | dtantsur | well, when something breaks, it's usually several attempts in a raw | |
| 14:33:41 | dansmith | either way, some documentation on "this may not be the behavior you want for ironic" is fine, but I definitely don't want to just say "for ironic this should be zero" | |
| 14:34:01 | dtantsur | I would be less opinionated here, but I suspect the users are going to see "no valid hosts found" in response, right? | |
| 14:34:15 | dansmith | once all the computes are disabled? yes. | |
| 14:34:17 | cdent | lajoskatona: I’ve responded on that review with some info on what’s going wrong. | |
| 14:34:53 | dansmith | dtantsur: the other thing that helps is that if you're really at 100% fail, eventually you stop wasting time, cpu, and network bandwidth trying to build and prepare things that are going to fail | |
| 14:34:57 | dtantsur | I'm reserving my opinion on this error :) it's too hard to debug, and it can mean anything (especially when RetryFilter is used) | |
| 14:35:18 | dansmith | so getting to zero disabled computes doesn't seem like a bad thing to me if literally 100% of the builds will fail | |
| 14:36:17 | dtantsur | does it include attempts from the RetryFilter? | |
| 14:36:51 | dansmith | I'm not sure what you mean.. are you asking if three retries against the same compute count as three failures? | |
| 14:37:05 | dtantsur | (side note: I became aware of this option after seeing https://review.openstack.org/#/c/496851/ ) | |
| 14:37:11 | dtantsur | dansmith: yep | |
| 14:37:23 | dansmith | dtantsur: sure, it's all the same | |
| 14:38:07 | dtantsur | dansmith: meaning, if people have retry count >= 10, one misconfigured instance can bring down a compute node? | |
| 14:38:29 | dansmith | dtantsur: if they configure retries high and don't adjust this, then sure | |
| 14:39:13 | dtantsur | okay, I think this is at least worth mentioning in our docs, because it may cause surprises (well, it did cause surprises already for tripleo people) | |
| 14:39:46 | openstackgerrit | Ed Leafe proposed openstack/nova master: docs: Document the scheduler workflow https://review.openstack.org/475810 | |
| 14:39:59 | dansmith | dtantsur: what sort of mention do you want? | |
| 14:40:26 | dansmith | note that the text of that config option mentions retries, | |
| 14:40:34 | dansmith | so hopefully someone upping that to a crazy level will see that | |
| 14:41:08 | dtantsur | dansmith: I want to have a mention of it in https://docs.openstack.org/ironic/latest/install/configure-compute.html | |
| 14:41:32 | dtantsur | maybe just something like "Mind the following option, if you don't like <behavior>, change it to 0" | |
| 14:42:15 | dansmith | dtantsur: sure, ironic docs that talk about ideal nova settings when using the two together is certainly a good idea :) | |
| 14:42:43 | dtantsur | yep, that was my initial idea | |
| 14:42:52 | dansmith | I have no problem with that | |
| 14:43:50 | alex_xu | mriedem: would you mind take care of https://review.openstack.org/496995, since the time is late for me | |
| 14:44:34 | dtantsur | cool, thanks all! | |
| 14:45:20 | mriedem | alex_xu: yup, i'll take over, thanks for working on it | |
| 14:45:47 | alex_xu | mriedem: thanks, good luck for rc2 :) | |
| 14:45:49 | mriedem | mnestratov|2: fyi https://bugs.launchpad.net/nova/+bug/1712801 | |
| 14:45:50 | openstack | Launchpad bug 1712801 in OpenStack Compute (nova) "Virtuozzo conainers lack SCSI info in libvirt XML" [Undecided,New] | |
| 14:46:36 | dtantsur | sahid: actually, could you take a quick look at https://docs.openstack.org/ironic/latest/install/configure-compute.html to check if everything is still correct there? there are a few options that I don't remember well | |
| 14:46:40 | dtantsur | oops, wrong ping | |
| 14:46:52 | dtantsur | sorry sahid, I meant dansmith (and please don't ask how I made this typo) | |
| 14:47:08 | sahid | :) | |
| 14:47:55 | mnestratov|2 | mriedem: thanks, this is our guy filed the bug in context of review https://review.openstack.org/#/c/495756/ | |
| 14:49:06 | dansmith | dtantsur: we don't need the hostmanager thing after the resource class transition right? | |
| 14:50:21 | dansmith | I'm also not sure why you tell people to use a crazy host_subset_size without explaining why they may or may not want that | |
| 14:50:36 | dansmith | especially in pike where we shouldn't have any scheduler races | |
| 14:51:12 | dtantsur | I think using a value > 1 is still a good idea, as two instances cannot get on one host | |
| 14:51:28 | dtantsur | probably not so huge, and I agree that we need an explanation | |
| 14:51:52 | dtantsur | and yes, it seems like we should have deprecated ironic_host_manager in pike | |
| 14:52:53 | lajoskatona | cdent: thanks I check it | |
| 14:53:06 | dansmith | dtantsur: with claims in the scheduler and resource classes, you can not have two instances pick the same host | |
| 14:53:18 | efried | edleafe "Return a list of selected host + alternates, along with their allocations to the conductor" <== Is [allocations to the conductor] a thing, or are [a list of selected hosts + alternates, along with their allocations] being returned to the conductor? | |
| 14:53:43 | efried | (Trying to figure out if there should be a comma after 'allocations'.) | |
| 14:53:49 | dtantsur | dansmith: right, I may be severely outdated on this | |
| 14:53:49 | dansmith | dtantsur: but without that, an insane host_subset_size just means there is zero weight applied to any host, which I would expect some people would not want | |
| 14:54:20 | edleafe | efried: the latter | |
| 14:54:21 | dtantsur | so, what would you recommend, keeping it 1? or setting it to something moderate? | |
| 14:54:35 | edleafe | we return a whole big glob of stuff to the conductor | |
| 14:54:45 | efried | edleafe k, so yes comma. There's another typo; will post an edit. | |
| 14:55:09 | dansmith | dtantsur: I would rather you explain why you're prescribing one thing or the other, and maybe just call out that it's moot with RC usage | |
| 14:55:38 | dtantsur | dansmith: then I wonder why we had to do https://review.openstack.org/#/c/493989/.. maybe it was before we moved the CI to resource classes though.. | |
| 14:55:41 | openstackgerrit | Eric Fried proposed openstack/nova master: docs: Document the scheduler workflow https://review.openstack.org/475810 | |
| 14:55:42 | dtantsur | vdrok: do you remember ^^^? | |
| 14:56:35 | dtantsur | or maybe it was because of backports... | |
| 14:56:37 | dansmith | dtantsur: yeah I dunno | |
| 14:56:49 | openstackgerrit | Eric Fried proposed openstack/nova master: Monkey patch the blockdiag extension https://review.openstack.org/476159 | |
| 14:56:55 | dtantsur | I cannot combine my brain back in a working state after the release drill | |
| 14:58:08 | dansmith | dtantsur: for devstack there's likely no reason to weigh any of the options and so increasing that value just means we make the scheduler choose randomly amongst empty hosts | |
| 14:58:15 | dansmith | which might be better for the gate for some reason, I dunno | |
| 14:58:33 | dansmith | but not for regular people that might want to weigh hosts with newer hardware differently or something like that | |
| 14:59:04 | dtantsur | weighing hosts for ironic is something new to me, but it makes sense indeed | |
| 14:59:33 | dansmith | dtantsur: well, maybe it's new because you disable weighing entirely with that 999999 thing :P | |
| 14:59:34 | edleafe | efried: I can't vim | |
| 14:59:44 | efried | edleafe :) | |
| 15:00:02 | dtantsur | dansmith: haha, no, I did not even know it was possible. have I ever mentioned my poor knowledge of nova? | |
| 15:00:12 | dansmith | dtantsur: I don't know that our in-tree weighers would mean anything for ironic nodes, but weighers are some of the more likely things for someone in a deployment to write customly to get the host selection they want | |
| 15:00:18 | dansmith | heh | |
| 15:00:42 | dtantsur | yeah, and "prefer newer hardware" makes some sense to me indeed, for example | |
| 15:00:58 | dansmith | yep | |
| 15:01:19 | openstackgerrit | sahid proposed openstack/nova master: libvirt: bandwidth param should be set in guest migrate https://review.openstack.org/497455 | |
| 15:01:19 | openstackgerrit | sahid proposed openstack/nova master: libvirt: add method to configure migration speed https://review.openstack.org/497456 | |
| 15:01:20 | openstackgerrit | sahid proposed openstack/nova master: libvirt: slowly live-migration to ensure network is ready https://review.openstack.org/497457 | |
| 15:01:27 | dtantsur | okay, thanks again dansmith. may I ping you for review when I write an update for this docs page? | |
| 15:01:38 | dansmith | dtantsur: sure | |
| 15:01:50 | dtantsur | cool, I'll try to finish it today | |
| 15:02:00 | openstackgerrit | David Rabel proposed openstack/nova master: Adds support for graceful shutdown for VMware instances https://review.openstack.org/494169 | |
| 15:06:21 | mriedem | here is the bug for the resize + reschedule not removing allocations issue https://bugs.launchpad.net/nova/+bug/1712850 | |
| 15:06:22 | openstack | Launchpad bug 1712850 in OpenStack Compute (nova) "Allocations are not removed from destination node when rescheduling during resize/migrate" [High,Triaged] | |
| 15:07:08 | vdrok | dtantsur: /me is on holiday today, but it was needed because all weights are the same in ci for our nodes, so the same ones were selected during parallel tests for instance build, and were racing as iirc claim happens on compute | |
| 15:07:51 | vdrok | So reschedules were happening constantly, and Max reschedules is 3 by default | |
| 15:08:04 | rabel | is "Intel PCI CI" always non-voting? | |
| 15:09:13 | rabel | that is: can i ignore that it failed without a reason? according to https://wiki.openstack.org/wiki/ThirdPartySystems/Intel-PCI-CI it is non-voting | |
| 15:09:33 | cdent | rabel: pretty much, yeah | |
| 15:09:49 | rabel | cool, thanks. | |
| 15:10:18 | mriedem | rabel: you can yell at sean-k-mooney about that | |
| 15:14:27 | dtantsur | dansmith: see vdrok's comment above ^^^. is it going away with resource classes? especially wrt "claim happens on compute"? | |