Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-10
14:04:02 sdague the important point is they met preconditions
14:04:06 sdague the service is fenced
14:04:12 sdague we don't care how they met those
14:04:19 sdague but they are expressing to nova that they did
14:04:24 bauzas sdague: we have a service group API for that
14:04:29 sdague it might have been a tool
14:05:01 bauzas sdague: the only usecase I heard of was that the SG API was lacking of functionality and either lagging or totally missing the host being down
14:05:21 bauzas sdague: so, eventually, the SG API would meet those preconds
14:05:49 bauzas that's just because Nova isn't intented to be a Nagios system, we allow other tools to fence the host for us
14:05:50 sdague bauzas: that's not good enough if I need it now
14:06:03 bauzas from our behalf, I mean
14:06:24 dtantsur edleafe: hi! my last attempt to use resource classes in the CI ended up with RamFilter removing the nodes
14:06:58 dtantsur I wonder if I'm missing something.. I thought I disabled requesting RAM/disk/CPU
14:07:50 bauzas dtantsur: Nova by default was still running those filters until yesterday
14:08:03 dtantsur oh
14:08:18 edleafe dtantsur: probably the request was for 1 of the resource class, along with the disk/ram/cpu in the flavor
14:08:36 dtantsur edleafe: I assume I'm removing the request for disk/ram/cpu from flavor
14:08:39 dtantsur lemme get a link
14:09:06 dtantsur edleafe: https://review.openstack.org/#/c/476968/13/devstack/lib/ironic@1889
14:09:41 dtantsur bauzas: this one, right? https://github.com/openstack/nova/commit/2fe96819c24eff5a9493a6559f3e8d5b4624a8c9
14:09:42 edleafe dtantsur: ok, then that should work
14:09:58 bauzas dtantsur: correct
14:10:10 dtantsur thanks, I'll see how it looks nowadays
14:10:10 edleafe dtantsur: do you have the call to placement anywhere in logs?
14:10:14 bauzas dtantsur: oh wait
14:10:23 bauzas dtantsur: Ironic is special-case IIRC
14:10:32 bauzas dtantsur: you folks have your own czay filters list :)
14:10:36 bauzas crazy
14:10:43 bauzas in nova
14:10:47 dtantsur yeah, we did change something around it.. lemme check
14:11:11 dtantsur https://review.openstack.org/#/c/490459/
14:12:07 dtantsur this is why I'm seeing the RamFilter, not the ExactRamFilter
14:12:20 dtantsur oh, and by the way. should we kill Exact filters with fire
14:12:25 dtantsur (well, I meant deprecate)
14:12:27 dtantsur ?
14:13:03 edleafe dtantsur: those were there only to support the pretense that an ironic node was a vm
14:13:11 edleafe dtantsur: so yeah, kill 'em!
14:13:20 dtantsur edleafe: wanna get a deprecation patch or should I?
14:13:45 edleafe dtantsur: we should also deprecate the separate ironic filter options, no?
14:14:05 dtantsur edleafe: yep
14:14:22 edleafe dtantsur: I may have time, but not much
14:14:38 edleafe dtantsur: if you want to start and post a WIP, I can pick it up
14:14:41 dtantsur ENOTMUCHTIME is a common error code nowadays
14:14:44 dtantsur sure, will do
14:20:19 bauzas dtantsur: those Exact* filters could just be treated like the other legacy filters
14:20:32 dtantsur and how do you treat legacy filters? :)
14:20:39 bauzas dtantsur: being removed from the list of filters to run by default, but still in tree for upgrade concerns
14:21:01 bauzas when you upgrade from Ocata, you certainly don't want to update nova.conf for that
14:21:11 dtantsur they're not on by default, there is an option to enable them..
14:21:30 openstackgerrit Balazs Gibizer proposed openstack/nova master: test server evacuation with placement https://review.openstack.org/492548
14:22:10 dtantsur folks, what's your next version? (to use with deprecated_since)
14:22:22 bauzas well, in theory you could use CachingScheduler with IroncHostManager I guess
14:22:31 bauzas in that case, you'd still require Exact* filters
14:22:43 gibi mriedem: just out of curiosity create a Pike -> Pike evac test and it seems the allocation on the source host has never cleaned up https://review.openstack.org/#/c/492548/
14:22:54 bauzas yet another call for deprecating the other scheduler driver we have in tree
14:23:21 dtantsur bauzas: I've never heard of people using it, but yeah. For every crazy feature there are people to try it in production..
14:23:26 ioggstream does anybody knows if soft-anti-affinity may be enabled in newton ?
14:23:38 bauzas ioggstream: IIRC, yes
14:24:15 bauzas ioggstream: https://blueprints.launchpad.net/nova/+spec/soft-affinity-for-server-group is Mitaka complete
14:24:16 ioggstream bauzas: by default it doesn't work but I saw that mitaka has an ERRATA
14:24:17 mriedem gibi: because the periodic task doesn't cleanup allocations anymore
14:24:24 mriedem gibi: it assumes the scheduler has everything correct
14:24:34 mriedem and the source node is 'down'
14:24:35 gibi mriedem: but not even the source compute cleans up?
14:24:46 gibi mriedem: after started up again?
14:24:48 mriedem if it's down we probably don't care about it
14:24:55 mriedem oh, we'll talk after the meeting
14:24:59 gibi mriedem: sure
14:26:15 mriedem but yeah the update_available_resource code in pike now does'nt overwrite the allocatoins
14:26:17 mriedem per that change
14:26:25 mriedem so that's why the source compute won't cleanup once it comes back up
14:26:32 mriedem it should remove though....
14:27:03 mriedem gibi: this one https://review.openstack.org/#/c/491850/
14:28:49 gibi mriedem: my test https://review.openstack.org/#/c/492548/ is top of https://review.openstack.org/#/c/491850/ and I still see the allocation on the source host
14:35:52 mriedem gibi: i think that's probably due to https://review.openstack.org/#/c/491012/12/nova/compute/resource_tracker.py@1047
14:36:19 gibi mriedem: checking the debug log...
14:36:38 mriedem gibi: but we should get into https://review.openstack.org/#/c/491012/12/nova/compute/resource_tracker.py@1145
14:36:44 mriedem _remove_deleted_instances_allocations
14:37:31 gibi mriedem: I see the debug log you pointed at
14:38:01 mriedem oh it could be https://review.openstack.org/#/c/491012/12/nova/compute/resource_tracker.py@1187
14:38:07 mriedem instance.node == cn.hypervisor_hostname):
14:38:07 mriedem if (instance.host == cn.host and
14:38:13 mriedem we continue there
14:38:19 mriedem or if instance.host != cn.host:
14:38:21 mriedem we also continue there
14:38:46 mriedem seems we should check to see if the instance is in self.tracked_migrations
14:38:47 gibi I can insert some extra log to confirm
14:40:49 gibi ahh there is logs already
14:41:12 gibi it is the instance.host == cn.host where we continue
14:43:17 jaypipes mriedem: ty
14:43:28 jaypipes mriedem: fyi, kinda vacationing today...
14:43:48 jaypipes mriedem: will work on my patches thouhg
14:45:36 mriedem jaypipes: don't think you have anything to work on
14:45:43 mriedem except follow ups for additional testing and whatnot
14:46:21 mriedem gibi: ok so self.tracked_migrations probably won't help after we restart the compute service since that dict will probably be empty
14:46:28 mriedem gibi: i left some comments in your test change,
14:46:34 mriedem we could maybe do some allocation cleanup in https://github.com/openstack/nova/blob/9a66d039a14afd591f4a3b6e655580aeeed17d29/nova/compute/manager.py#L649
14:46:48 jaypipes mriedem: yeah
14:47:24 mriedem gibi: this is similar to https://bugs.launchpad.net/nova/+bug/1679750 where we don't delete the allocations for the instance on the compute host during a 'local delete' in the API
14:47:25 openstack Launchpad bug 1679750 in OpenStack Compute (nova) "Allocations are not cleaned up in placement for instance 'local delete' case" [Medium,In progress]

Earlier   Later