Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-10
13:20:28 cdent but probably good given what we want eventually
13:21:25 sdague mriedem: sure but as you know, there aren't many idle folks looking for bugs to fix. Docs bugs mostly languish in our tracker.
13:21:36 mriedem gibi: ok so the doctor project is supposed to get an alarm that something is wrong with the host, fence it and then force it down and start evacuating?
13:21:49 mriedem sdague: i planned on fixing that docs gap myself
13:21:53 mriedem but $time
13:21:59 gibi mriedem: I think so, yes
13:22:02 sdague mriedem: sure, that's fine
13:22:03 mriedem what does 'fencing' mean in this case?
13:22:06 bauzas mriedem: do we have a devstack change for https://review.openstack.org/#/c/491854/ ?
13:22:16 cdent re: $time [t 1zqO]
13:22:16 purplerbot <cdent> When do we start asking if the concept of PTL, as currently constructed, is sustainable? [2017-08-10 13:21:20.802315] [n 1zqO]
13:22:20 mriedem bauzas: no
13:22:22 gibi mriedem: power off, or cut the network
13:22:32 gibi mriedem: mostly power off via IPMI
13:22:33 bauzas mriedem: okay, I'll write it
13:22:34 sdague what gibi said
13:22:48 mriedem ok powering off would be ideal
13:23:04 sdague mriedem: but, it could be lots of things. They could also decide to network fence the node
13:23:05 mriedem but network is also good so the compute couldn't send changes to the report client (to conductor i mean)
13:23:19 mriedem as long as it can't get to the placement api then that's sufficient
13:23:29 cdent that comment from bauzas reminds me of something I read while reading gerrit messages last night: did some test have to be changed so that it was running one of the filters we have no declared no longer default?
13:23:33 gibi mriedem: I'm not 100% sure doctor also starts the evacuation automatically after force_down
13:23:53 mriedem gibi: ok but some project, maybe watcher
13:23:58 gibi mriedem: sure
13:23:58 sdague mriedem: I can take a spin on the api-ref, if you review my other doc patches :)
13:24:03 gibi mriedem: we have our on internally :)
13:24:10 bauzas what's the problem with evacuation ?
13:24:19 mriedem sdague: i have one change i need to get in today and that's https://review.openstack.org/#/c/491012/
13:24:43 mriedem bauzas: see https://review.openstack.org/#/c/491012/
13:25:39 bauzas mriedem: ah, this...
13:29:05 mriedem cdent: ok left more comments in there
13:29:14 mriedem i think i'm ok with this now, or as much as i can be,
13:29:26 mriedem cdent: i'll fix the functional test quick and push that up, just that change
13:29:55 mriedem and then i'll probably spend some time this afternoon working on a functional test for evacuate
13:30:01 mriedem with the service version stuff
13:30:39 mriedem bauzas: devstack sets this
13:30:40 mriedem iniset $NOVA_CONF filter_scheduler enabled_filters "RetryFilter,AvailabilityZoneFilter,ComputeFilter,ComputeCapabilitiesFilter,ImagePropertiesFilter,CoreFilter,RamFilter,DiskFilter"
13:30:52 bauzas mriedem: I know
13:30:58 mriedem nvm, that's only for the fake hypervisor
13:30:59 bauzas mriedem: I'm just writing atm the change to remove this
13:31:11 mriedem which was only used in the large tests, which we don't run anymore
13:31:23 bauzas we run Core and Ram by default AFAIK
13:31:31 bauzas not Disk tho
13:31:53 mriedem FILTERS="RetryFilter,AvailabilityZoneFilter,RamFilter,DiskFilter,ComputeFilter,ComputeCapabilitiesFilter,ImagePropertiesFilter,ServerGroupAntiAffinityFilter,ServerGroupAffinityFilter,SameHostFilter,DifferentHostFilter"
13:32:08 mriedem ^ is what we run by default
13:32:41 bauzas so, we need to remove Ram *and* Diks
13:32:56 bauzas I was just remembering we still run 2 legacy filters
13:32:59 mriedem http://git.openstack.org/cgit/openstack-dev/devstack/tree/lib/nova#n104
13:33:02 mriedem ok
13:33:16 bauzas anyway, patch is on its way
13:34:07 mriedem sdague: oh right, that docs bug was because cfriesen was asking about this, see the last paragraph or two in that bug report,
13:34:30 mriedem someone had forced the service down, upgraded everything else, and then when they tried to set forced_down=False, it failed with ServiceTooOld
13:34:50 mriedem which i think just caused some confusion in how that all works
13:35:01 openstackgerrit Sean Dague proposed openstack/nova master: Update api-guide and api-ref to be clear about forced-down https://review.openstack.org/492533
13:35:07 mriedem the ServiceTooOld is i'm pretty sure a 500 out of the PUT /os-services API too
13:35:40 sdague mriedem: yeh, I don't know what the recovery path is on that all, but at least we can be clear on what the state transition into it should look like
13:35:46 cdent does anyone have anything specific for me to do, or shall I continue chasing reviews and randomly testing things by hand?
13:37:22 edleafe cdent: sudo make me a sandwich
13:38:50 bauzas sdague: mriedem: https://review.openstack.org/#/c/492537/
13:39:04 bauzas ^ devstack change removing legacy filters FTW
13:39:21 sdague bauzas: ok, cool
13:39:28 sdague bauzas: is there a reason we override those in the first place?
13:39:39 bauzas sdague: correct me if I'm wrong but we don't need to modify grenade since it uses default config optoions ?
13:39:41 sdague would we just use defaults?
13:39:44 mriedem sdague: because devstack had the default + SameHost + DifferentHost
13:39:49 bauzas what mriedem said
13:39:57 mriedem tempest has tests for SameHost/DifferentHost filters
13:39:59 sdague ah, cool, good reason
13:40:00 mriedem which aren't defaults
13:40:02 bauzas we could tho make += "SameHost"
13:40:09 bauzas since it's a listopt
13:40:15 mriedem in bash?
13:40:23 sdague bauzas: that doesn't work in setting nova.conf
13:40:23 bauzas ah, right
13:40:35 bauzas I usually do this directly :)
13:41:01 bauzas sdague: anyway, my question still remains wrt grenade
13:41:13 bauzas sdague: do we need to s// something when we upgrade the node ?
13:41:23 sdague bauzas: grenade should be fine as long as those things didn't get deleted
13:41:25 bauzas in terms of nova.conf -ism
13:41:36 sdague it will use the ocata config in pike
13:41:46 bauzas yeah, my thoughts
13:41:48 bauzas not super then
13:42:13 bauzas we don't remove those filters because CachingScheduler still requires them
13:42:28 bauzas but in theory, you don't need to run those filters since Ocata
13:42:40 sdague right as long as there is normal deprecation cycles you don't need to do anything in grenade
13:42:51 bauzas that's just we wanted to be gentle in Ocata and not ask operators to change their conf when upgrading from Newton
13:42:59 sdague as long as you remove the deprecated setting of things in devstack in the release where you deprecate them
13:43:08 bauzas okay, I think it's reasonable then
13:43:10 sdague it just naturally rolls through it all
13:43:17 bauzas k
13:44:33 bauzas mriedem: looping over https://etherpad.openstack.org/p/nova-pike-release-candidate-todo
13:44:48 sdague mriedem: I think putting the entire recovery path inside the API ref for forced-down is not idea - https://review.openstack.org/#/c/492533/1/api-ref/source/os-services.inc
13:45:02 bauzas mriedem: AFAICT, the top prio for reviewing is then https://review.openstack.org/#/c/491012/ ?
13:45:06 sdague there probably needs to be a specific HA recovery document for the whole thing
13:45:15 sdague this is just "don't use this unless you really are sure"
13:45:34 bauzas it's a force thing
13:45:34 bauzas :p
13:45:51 bauzas 'force' is meaningful in Linux terminology :)
13:46:08 mriedem bauzas: yes
13:46:15 bauzas don't expect good things to happen magically if you use a force flag

Earlier   Later