Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-10
13:02:41 mriedem artom: you love evacuate right?
13:02:54 mriedem gibi: you love finding bugs right?
13:05:47 gibi mriedem: I would put it I like finding them now than getting it from production :)
13:05:52 artom mriedem, in the same way I love, err...
13:06:01 artom Crap, it's too early for witty wordplay
13:06:08 artom mriedem, anyways, what's up?
13:06:22 mriedem my main worry evacuate from an ocata compute messing this up https://review.openstack.org/#/c/491012/
13:06:27 mriedem artom: i don't know how much you've followed this
13:06:37 mriedem but basically the filter scheduler creates allocations in placement now,
13:06:43 mriedem on both the source and dest computes during a move
13:06:45 mriedem like evacuate
13:07:28 mriedem the problem is that the resource tracker has no concept of other providers than itself, so during it's periodic accounting updates, it overwrites allocations in placement for any other provider
13:07:37 mriedem that patch ^ attempts to resolve that
13:08:03 mriedem by using a minimum compute service version check - so once all of the computes are pike, it will stop doing it's local accounting
13:08:10 mriedem and overwriting the stuff the scheduler created
13:08:54 mriedem one of my worries is that we have an ocata compute that is forced-down, which takes it out of the service version check, but could still be running and trampling on things
13:09:47 mriedem i think it's probably a small window because if you are forcing a compute down and evacuating from it, (1) you're likely to stop that host at some point and (2) once the instances move, the dest compute should be accounting for them - and the scheduler will also do that
13:10:43 artom I'm fuzzy on the resource tracker having no concept of other providers than itself
13:11:02 artom I thought compute nodes were resource providers?
13:11:21 mriedem they are
13:11:32 openstackgerrit Merged openstack/nova master: Keep the code consistent https://review.openstack.org/490304
13:12:04 mriedem cdent: am i correct in saying that https://review.openstack.org/#/c/491012/ only applies to the periodic update_available_resource task?
13:12:15 mriedem looks like that's the only place that _update_usage_from_instances is called from
13:12:19 openstackgerrit Merged openstack/nova master: add description about key_name https://review.openstack.org/489525
13:13:07 cdent mriedem: yes
13:13:10 mriedem just thinking that if that code thinks everything is pike and doesn't auto-heal,
13:13:23 mriedem and it misses the instance moving from the ocata compute,
13:13:31 mriedem then we have to be sure that the rebuild_claim handles it
13:14:44 cdent murgh
13:14:54 gibi mriedem: if a compute host is forced_down it should mean that that compute host is fenced
13:15:16 gibi mriedem: therefore it cannot tramp on allocations
13:15:50 mriedem gibi: that doesn't mean the nova-compute service is not running on that host
13:16:11 mriedem and if the service is running, it's update_available_resource periodic is running and could be overwriting allocations for the instance that's being evacuated
13:16:40 mriedem the forced_down flag doesn't do anything besides let the evacuate API proceed before the servicegroup api checkin says the compute is down
13:16:55 openstackgerrit Markus Zoeller (markus_z) proposed openstack/nova master: docs: Explain the flow of the "serial console" feature https://review.openstack.org/476188
13:17:19 mriedem cdent: so i don't think rebuild_claim will update allocations at all
13:17:31 sdague mriedem: the contract with the user is forced_down means they killed that compute
13:17:37 gibi mriedem: if the compute is stull running but the admin set force-down then it is a user error
13:17:49 gibi admin should fence first then set forced-down flag
13:17:51 mriedem rebuild_call calls _move_call which calls _update_usage_from_migration which calls _update_usage which doesn't call the report client
13:17:52 sdague it is only meant to be used if they've taken that system out of communication
13:18:15 sdague agree with gibi, that's admin error, and we've never attempted to correct for that
13:18:33 mriedem that's not documented anywhere https://developer.openstack.org/api-ref/compute/#update-forced-down
13:18:47 sdague mriedem: ok, well we should document it, that was the whole point of that feature
13:19:05 sdague for HA systems to override nova when it knew better
13:19:12 cdent mriedem: yeah, it looks like the allocation creation is all happening outside the various _claim* methods
13:19:15 mriedem i had reported a bug related to docs on this at one point https://bugs.launchpad.net/nova/+bug/1691871
13:19:16 openstack Launchpad bug 1691871 in OpenStack Compute (nova) "forced-down vs service disable is not documented well in the compute API reference" [Medium,Confirmed]
13:19:16 gibi mriedem: it is at least in the original spec https://specs.openstack.org/openstack/nova-specs/specs/liberty/implemented/mark-host-down.html
13:19:18 cdent which is somewhat weird
13:20:28 cdent but probably good given what we want eventually
13:21:25 sdague mriedem: sure but as you know, there aren't many idle folks looking for bugs to fix. Docs bugs mostly languish in our tracker.
13:21:36 mriedem gibi: ok so the doctor project is supposed to get an alarm that something is wrong with the host, fence it and then force it down and start evacuating?
13:21:49 mriedem sdague: i planned on fixing that docs gap myself
13:21:53 mriedem but $time
13:21:59 gibi mriedem: I think so, yes
13:22:02 sdague mriedem: sure, that's fine
13:22:03 mriedem what does 'fencing' mean in this case?
13:22:06 bauzas mriedem: do we have a devstack change for https://review.openstack.org/#/c/491854/ ?
13:22:16 cdent re: $time [t 1zqO]
13:22:16 purplerbot <cdent> When do we start asking if the concept of PTL, as currently constructed, is sustainable? [2017-08-10 13:21:20.802315] [n 1zqO]
13:22:20 mriedem bauzas: no
13:22:22 gibi mriedem: power off, or cut the network
13:22:32 gibi mriedem: mostly power off via IPMI
13:22:33 bauzas mriedem: okay, I'll write it
13:22:34 sdague what gibi said
13:22:48 mriedem ok powering off would be ideal
13:23:04 sdague mriedem: but, it could be lots of things. They could also decide to network fence the node
13:23:05 mriedem but network is also good so the compute couldn't send changes to the report client (to conductor i mean)
13:23:19 mriedem as long as it can't get to the placement api then that's sufficient
13:23:29 cdent that comment from bauzas reminds me of something I read while reading gerrit messages last night: did some test have to be changed so that it was running one of the filters we have no declared no longer default?
13:23:33 gibi mriedem: I'm not 100% sure doctor also starts the evacuation automatically after force_down
13:23:53 mriedem gibi: ok but some project, maybe watcher
13:23:58 gibi mriedem: sure
13:23:58 sdague mriedem: I can take a spin on the api-ref, if you review my other doc patches :)
13:24:03 gibi mriedem: we have our on internally :)
13:24:10 bauzas what's the problem with evacuation ?
13:24:19 mriedem sdague: i have one change i need to get in today and that's https://review.openstack.org/#/c/491012/
13:24:43 mriedem bauzas: see https://review.openstack.org/#/c/491012/
13:25:39 bauzas mriedem: ah, this...
13:29:05 mriedem cdent: ok left more comments in there
13:29:14 mriedem i think i'm ok with this now, or as much as i can be,
13:29:26 mriedem cdent: i'll fix the functional test quick and push that up, just that change
13:29:55 mriedem and then i'll probably spend some time this afternoon working on a functional test for evacuate
13:30:01 mriedem with the service version stuff
13:30:39 mriedem bauzas: devstack sets this
13:30:40 mriedem iniset $NOVA_CONF filter_scheduler enabled_filters "RetryFilter,AvailabilityZoneFilter,ComputeFilter,ComputeCapabilitiesFilter,ImagePropertiesFilter,CoreFilter,RamFilter,DiskFilter"
13:30:52 bauzas mriedem: I know
13:30:58 mriedem nvm, that's only for the fake hypervisor
13:30:59 bauzas mriedem: I'm just writing atm the change to remove this
13:31:11 mriedem which was only used in the large tests, which we don't run anymore
13:31:23 bauzas we run Core and Ram by default AFAIK
13:31:31 bauzas not Disk tho
13:31:53 mriedem FILTERS="RetryFilter,AvailabilityZoneFilter,RamFilter,DiskFilter,ComputeFilter,ComputeCapabilitiesFilter,ImagePropertiesFilter,ServerGroupAntiAffinityFilter,ServerGroupAffinityFilter,SameHostFilter,DifferentHostFilter"
13:32:08 mriedem ^ is what we run by default
13:32:41 bauzas so, we need to remove Ram *and* Diks
13:32:56 bauzas I was just remembering we still run 2 legacy filters
13:32:59 mriedem http://git.openstack.org/cgit/openstack-dev/devstack/tree/lib/nova#n104
13:33:02 mriedem ok
13:33:16 bauzas anyway, patch is on its way

Earlier   Later