Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-10
13:09:47 mriedem i think it's probably a small window because if you are forcing a compute down and evacuating from it, (1) you're likely to stop that host at some point and (2) once the instances move, the dest compute should be accounting for them - and the scheduler will also do that
13:10:43 artom I'm fuzzy on the resource tracker having no concept of other providers than itself
13:11:02 artom I thought compute nodes were resource providers?
13:11:21 mriedem they are
13:11:32 openstackgerrit Merged openstack/nova master: Keep the code consistent https://review.openstack.org/490304
13:12:04 mriedem cdent: am i correct in saying that https://review.openstack.org/#/c/491012/ only applies to the periodic update_available_resource task?
13:12:15 mriedem looks like that's the only place that _update_usage_from_instances is called from
13:12:19 openstackgerrit Merged openstack/nova master: add description about key_name https://review.openstack.org/489525
13:13:07 cdent mriedem: yes
13:13:10 mriedem just thinking that if that code thinks everything is pike and doesn't auto-heal,
13:13:23 mriedem and it misses the instance moving from the ocata compute,
13:13:31 mriedem then we have to be sure that the rebuild_claim handles it
13:14:44 cdent murgh
13:14:54 gibi mriedem: if a compute host is forced_down it should mean that that compute host is fenced
13:15:16 gibi mriedem: therefore it cannot tramp on allocations
13:15:50 mriedem gibi: that doesn't mean the nova-compute service is not running on that host
13:16:11 mriedem and if the service is running, it's update_available_resource periodic is running and could be overwriting allocations for the instance that's being evacuated
13:16:40 mriedem the forced_down flag doesn't do anything besides let the evacuate API proceed before the servicegroup api checkin says the compute is down
13:16:55 openstackgerrit Markus Zoeller (markus_z) proposed openstack/nova master: docs: Explain the flow of the "serial console" feature https://review.openstack.org/476188
13:17:19 mriedem cdent: so i don't think rebuild_claim will update allocations at all
13:17:31 sdague mriedem: the contract with the user is forced_down means they killed that compute
13:17:37 gibi mriedem: if the compute is stull running but the admin set force-down then it is a user error
13:17:49 gibi admin should fence first then set forced-down flag
13:17:51 mriedem rebuild_call calls _move_call which calls _update_usage_from_migration which calls _update_usage which doesn't call the report client
13:17:52 sdague it is only meant to be used if they've taken that system out of communication
13:18:15 sdague agree with gibi, that's admin error, and we've never attempted to correct for that
13:18:33 mriedem that's not documented anywhere https://developer.openstack.org/api-ref/compute/#update-forced-down
13:18:47 sdague mriedem: ok, well we should document it, that was the whole point of that feature
13:19:05 sdague for HA systems to override nova when it knew better
13:19:12 cdent mriedem: yeah, it looks like the allocation creation is all happening outside the various _claim* methods
13:19:15 mriedem i had reported a bug related to docs on this at one point https://bugs.launchpad.net/nova/+bug/1691871
13:19:16 openstack Launchpad bug 1691871 in OpenStack Compute (nova) "forced-down vs service disable is not documented well in the compute API reference" [Medium,Confirmed]
13:19:16 gibi mriedem: it is at least in the original spec https://specs.openstack.org/openstack/nova-specs/specs/liberty/implemented/mark-host-down.html
13:19:18 cdent which is somewhat weird
13:20:28 cdent but probably good given what we want eventually
13:21:25 sdague mriedem: sure but as you know, there aren't many idle folks looking for bugs to fix. Docs bugs mostly languish in our tracker.
13:21:36 mriedem gibi: ok so the doctor project is supposed to get an alarm that something is wrong with the host, fence it and then force it down and start evacuating?
13:21:49 mriedem sdague: i planned on fixing that docs gap myself
13:21:53 mriedem but $time
13:21:59 gibi mriedem: I think so, yes
13:22:02 sdague mriedem: sure, that's fine
13:22:03 mriedem what does 'fencing' mean in this case?
13:22:06 bauzas mriedem: do we have a devstack change for https://review.openstack.org/#/c/491854/ ?
13:22:16 cdent re: $time [t 1zqO]
13:22:16 purplerbot <cdent> When do we start asking if the concept of PTL, as currently constructed, is sustainable? [2017-08-10 13:21:20.802315] [n 1zqO]
13:22:20 mriedem bauzas: no
13:22:22 gibi mriedem: power off, or cut the network
13:22:32 gibi mriedem: mostly power off via IPMI
13:22:33 bauzas mriedem: okay, I'll write it
13:22:34 sdague what gibi said
13:22:48 mriedem ok powering off would be ideal
13:23:04 sdague mriedem: but, it could be lots of things. They could also decide to network fence the node
13:23:05 mriedem but network is also good so the compute couldn't send changes to the report client (to conductor i mean)
13:23:19 mriedem as long as it can't get to the placement api then that's sufficient
13:23:29 cdent that comment from bauzas reminds me of something I read while reading gerrit messages last night: did some test have to be changed so that it was running one of the filters we have no declared no longer default?
13:23:33 gibi mriedem: I'm not 100% sure doctor also starts the evacuation automatically after force_down
13:23:53 mriedem gibi: ok but some project, maybe watcher
13:23:58 gibi mriedem: sure
13:23:58 sdague mriedem: I can take a spin on the api-ref, if you review my other doc patches :)
13:24:03 gibi mriedem: we have our on internally :)
13:24:10 bauzas what's the problem with evacuation ?
13:24:19 mriedem sdague: i have one change i need to get in today and that's https://review.openstack.org/#/c/491012/
13:24:43 mriedem bauzas: see https://review.openstack.org/#/c/491012/
13:25:39 bauzas mriedem: ah, this...
13:29:05 mriedem cdent: ok left more comments in there
13:29:14 mriedem i think i'm ok with this now, or as much as i can be,
13:29:26 mriedem cdent: i'll fix the functional test quick and push that up, just that change
13:29:55 mriedem and then i'll probably spend some time this afternoon working on a functional test for evacuate
13:30:01 mriedem with the service version stuff
13:30:39 mriedem bauzas: devstack sets this
13:30:40 mriedem iniset $NOVA_CONF filter_scheduler enabled_filters "RetryFilter,AvailabilityZoneFilter,ComputeFilter,ComputeCapabilitiesFilter,ImagePropertiesFilter,CoreFilter,RamFilter,DiskFilter"
13:30:52 bauzas mriedem: I know
13:30:58 mriedem nvm, that's only for the fake hypervisor
13:30:59 bauzas mriedem: I'm just writing atm the change to remove this
13:31:11 mriedem which was only used in the large tests, which we don't run anymore
13:31:23 bauzas we run Core and Ram by default AFAIK
13:31:31 bauzas not Disk tho
13:31:53 mriedem FILTERS="RetryFilter,AvailabilityZoneFilter,RamFilter,DiskFilter,ComputeFilter,ComputeCapabilitiesFilter,ImagePropertiesFilter,ServerGroupAntiAffinityFilter,ServerGroupAffinityFilter,SameHostFilter,DifferentHostFilter"
13:32:08 mriedem ^ is what we run by default
13:32:41 bauzas so, we need to remove Ram *and* Diks
13:32:56 bauzas I was just remembering we still run 2 legacy filters
13:32:59 mriedem http://git.openstack.org/cgit/openstack-dev/devstack/tree/lib/nova#n104
13:33:02 mriedem ok
13:33:16 bauzas anyway, patch is on its way
13:34:07 mriedem sdague: oh right, that docs bug was because cfriesen was asking about this, see the last paragraph or two in that bug report,
13:34:30 mriedem someone had forced the service down, upgraded everything else, and then when they tried to set forced_down=False, it failed with ServiceTooOld
13:34:50 mriedem which i think just caused some confusion in how that all works
13:35:01 openstackgerrit Sean Dague proposed openstack/nova master: Update api-guide and api-ref to be clear about forced-down https://review.openstack.org/492533
13:35:07 mriedem the ServiceTooOld is i'm pretty sure a 500 out of the PUT /os-services API too
13:35:40 sdague mriedem: yeh, I don't know what the recovery path is on that all, but at least we can be clear on what the state transition into it should look like
13:35:46 cdent does anyone have anything specific for me to do, or shall I continue chasing reviews and randomly testing things by hand?
13:37:22 edleafe cdent: sudo make me a sandwich
13:38:50 bauzas sdague: mriedem: https://review.openstack.org/#/c/492537/
13:39:04 bauzas ^ devstack change removing legacy filters FTW
13:39:21 sdague bauzas: ok, cool
13:39:28 sdague bauzas: is there a reason we override those in the first place?
13:39:39 bauzas sdague: correct me if I'm wrong but we don't need to modify grenade since it uses default config optoions ?
13:39:41 sdague would we just use defaults?
13:39:44 mriedem sdague: because devstack had the default + SameHost + DifferentHost
13:39:49 bauzas what mriedem said

Earlier   Later