Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-24
13:02:25 openstackgerrit Gábor Antal proposed openstack/nova master: Transform instance.live_migration_rollback_dest notification https://review.openstack.org/480214
13:07:46 VAhl Is it possible to change port on the Placement-api? If yes, which parameter should be set in the [placement] section in nova.conf
13:10:29 VAhl found it. It is sites-available/nova-placement-api.conf
13:11:48 stephenfin mriedem: I cherry-picked that cell deprecation to stable/pike, if we want to bring it in https://review.openstack.org/#/c/497166
13:11:59 stephenfin (mostly so we can stay compliant with the deprecation policy)
13:12:58 mriedem :/
13:22:56 openstackgerrit Eric Fried proposed openstack/nova master: Monkey patch the blockdiag extension https://review.openstack.org/476159
13:23:04 mriedem alex_xu: thanks for starting this https://review.openstack.org/#/c/496995/ - i've got some comments inline
13:23:09 mriedem cdent: dansmith: ^
13:30:22 rabel hi there. we have a patch for the vmware driver: https://review.openstack.org/#/c/494169/ ( cdent already reviewed it ). i posted a comment with vmware-recheck-patch about a week ago, but nothing happens. can you help me on how to trigger vmware ci tests?
13:32:24 cdent rabel: you may have a hit time when the vmware ci wasn’t happy. I’ve tried to trigger another
13:32:34 rabel cdent: thank you!
13:32:43 mriedem it hasn't been happy for about 6 months
13:33:04 cdent mriedem: I’m happy to report that that is almost but not quite on the agenda to get some attention soon
13:33:04 mriedem but, there haven't been a ton of changes proposed to make it a priority
13:33:20 mriedem heh
13:34:06 mriedem alex_xu: i'm not sure if you want to handle the reschedule for resize in the same patch, i know it's getting late for you and we don't have a functional regression test for the resize + reschedule case
13:34:11 cdent there’s been the usual such and such is in the way
13:38:59 alex_xu mriedem: yea, I also thought the resize case is another patch
13:39:26 alex_xu mriedem: cdent just replied this comment https://review.openstack.org/#/c/496995/5/nova/compute/manager.py@1733
13:41:13 lajoskatona cdent: Gibi asked me to think about adding an extra testclass to test_servers to test the server moving stuff with custom resources
13:41:56 lajoskatona cdent: now I have a new class, and the setup works with custom resources and resorce providers, whatever....
13:42:05 cdent lajoskatona: that’s a good idea
13:43:24 lajoskatona cdent: BUT: all tests are failing, due to exception from scheduler: 2017-08-24 14:49:41,491 DEBUG [nova.scheduler.manager] Got no allocation candidates from the Placement API. This may be a temporary occurrence as compute nodes start up and begin reporting inventory to the Placement service.
13:44:03 lajoskatona cdent: shall I upload as WIP, and somebody should check it what am I missing from the workflow?
13:44:03 cdent can you push up your code as a wip so it is easier to look at?
13:44:06 cdent :)
13:44:11 cdent yes!
13:44:16 lajoskatona cdent, ok, I upload :-)
13:44:23 alex_xu mriedem: cdent in the original RT behaviour, the instance claim will release the resource, we also nill out instance's host and name, that also means the resource will be released by RT
13:45:32 mriedem i'm replying
13:47:51 mriedem guh
13:47:54 mriedem ok replied
13:48:07 mriedem alex_xu: so you've got a point, as usual :)
13:48:28 mriedem if the instance.host is None, then the API isn't going to cast to the compute to complete the deletion which would remove the allocation
13:48:39 mriedem and we don't currently remove the allocations in the API during a 'local' delete
13:48:45 mriedem so we're kind of stuck
13:49:02 cdent is there a reason we don’t do it from the API?
13:49:09 cdent s/don’t/don’t want to/
13:49:09 mriedem hasn't been fixed yet
13:49:17 openstack Launchpad bug 1679750 in OpenStack Compute (nova) "Allocations are not cleaned up in placement for instance 'local delete' case" [Medium,Confirmed]
13:49:17 mriedem no, it's just an open bug https://bugs.launchpad.net/nova/+bug/1679750
13:49:37 mriedem cdent: it was less of a problem at the time it was reported because the periodic task would heal and remove the allocations
13:50:02 mriedem well, actually let me see about that
13:50:41 mriedem yeah, actually the periodic task should still clean these up, in _remove_deleted_instances_allocations
13:50:46 mriedem that gets all allocations for the node
13:51:09 mriedem and if we get InstanceNotFound looking up the instance, it deletes the allocations for that instance
13:55:06 mriedem left some more comments
13:56:05 mriedem probably need to get dansmith's opinion when he's around. i think we could go with this patch even though the periodic would cleanup the allocations for the node once the instance is deleted, and if we removed the allocations from the api during local delete that would also take care of it - but we don't have that fix yet
13:57:08 dansmith about whether or not we should delete allocations on local delete/
13:59:43 mriedem no,
13:59:47 mriedem https://review.openstack.org/#/c/496995/
14:00:04 mriedem on build_results.FAILED - so unexpected failures during build, or BuildAbortException
14:00:09 mriedem but, it's meeting time
14:08:20 openstackgerrit Lajos Katona proposed openstack/nova master: WIP: Test server movings with custom resources https://review.openstack.org/497399
14:10:32 cdent lajoskatona: will take a look at that ^ once the nova meeting is done
14:15:25 openstackgerrit Eric Fried proposed openstack/nova master: Add formatting to scheduling activity diagram https://review.openstack.org/476204
14:17:35 efried edleafe stephenfin sfinucan ^^
14:17:50 stephenfin efried: Cheers. Looking
14:18:20 efried stephenfin edleafe Rebased on top of the monkeypatch one and updated for the latest text in the doc.
14:18:30 efried Not too much to see until the docs build completes, I guess :)
14:20:08 openstackgerrit Merged openstack/nova master: How about not logging errors every time we shelve offload? https://review.openstack.org/496930
14:20:11 stephenfin efried: Yup, that's what I was hoping for. +2d, on the assumption the gate will catch any typos :)
14:20:56 dtantsur hi folks, it's me again :) do you think we should document setting https://github.com/openstack/nova/blob/master/nova/conf/compute.py#L634 to 0 in case of ironic?
14:21:14 dtantsur I don't think the compute instance is a popular cause of build failures for ironic, as opposed to the even world around it
14:21:58 efried dtantsur Are you saying ironic overrides that default today, or are you suggesting that it do so?
14:22:19 dtantsur efried: I'm suggesting we document overriding it for nova-compute processes bound to ironic
14:22:30 dtantsur (maybe in ironic docs, we have a section on configuring nova)
14:22:38 dtantsur I'm just checking if my understanding of it is correct
14:22:57 efried dtantsur Oh, you're suggesting no code changes, just documenting that the user, if using ironic, will want to set it to zero.
14:23:07 dtantsur correct
14:24:01 efried dtantsur So the way I understand that config var, it'll disable the compute host entirely if <N> consecutive builds fail. On the theory that we ought to stop trying to schedule to that host, because it's obviously sick. Why is ironic different in this regard?
14:24:40 dtantsur efried: because the nova-compute hosts merely pipe requests to ironic
14:25:12 dtantsur disabling a compute host will make some part of your ironic fleet unavailable, and that may NOT be the part that caused failures
14:26:00 efried I would think that's always the case ("may NOT be the part that caused failures").
14:26:20 openstackgerrit Alex Xu proposed openstack/nova master: Remove allocation when booting instance rescheduled or aborted https://review.openstack.org/496995
14:26:23 dtantsur yeah, usually our failures are caused by network problems or misconfiguration of ironic
14:26:31 efried (btw dtantsur I'm not arguing against what you're suggesting - just playing devil's advocate)
14:26:31 dtantsur not by the nova-compute process being broken
14:26:42 dtantsur yeah, got it :)
14:27:04 efried Well, right; if it's network problems that are persistent enough to cause 10 consecutive failures, you should go fix the problems and then re-enable the compute host.
14:27:12 efried Ditto with misconfiguration, even more so.
14:27:35 dtantsur then why kill nova-compute, but not nova-api? :)
14:27:50 dtantsur both are not the source of the problem, both merely pass the requests on
14:28:14 efried Heh, guess that's a dansmith question.
14:28:15 dtantsur also, not the whole ironic deployment may be broken, but you're disabling a *random* part of it
14:28:42 efried That doesn't seem like an argument specific to ironic.
14:29:02 dtantsur no, because nova-compute for e.g. libvirt is tied to its libvirt instances
14:29:03 dansmith so,
14:29:11 dtantsur while nova-compute for ironic handles some randomly chosen share of ironic nodes
14:29:16 dtantsur which can change with time, btw
14:29:20 dansmith the request for that came with ironic people in the room, talking about such failures, and I'm pretty sure they were on board
14:29:50 dansmith dtantsur: if you fail to build ten things _consecutively_ don't you think that's cause for concern?
14:29:56 dtantsur can we can specific people please then? I'm just seeing it as a source of wrong bug reports coming to me
14:30:07 dansmith like, why not drop that compute out and let the ring rebalance?
14:30:24 dtantsur dansmith: rebalanace annnnndd? fail 10 times more, because the problem was not fixed?
14:30:31 dtantsur and this way until we run out of computes?
14:30:44 dansmith dtantsur: well, it depends on what the problem is of course
14:31:15 dtantsur my point is that with VMs killing a nova-compute process can cut a misbahaving compute node. with ironic it does mostly nothing, until we run out of computes (that's where everything breaks)
14:31:27 dansmith dtantsur: the conversation in the room centered around (a) not letting one compute node be a black hole and (b) making it super obvious that something was wrong

Earlier   Later