| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-10 | |||
| 13:16:11 | mriedem | and if the service is running, it's update_available_resource periodic is running and could be overwriting allocations for the instance that's being evacuated | |
| 13:16:40 | mriedem | the forced_down flag doesn't do anything besides let the evacuate API proceed before the servicegroup api checkin says the compute is down | |
| 13:16:55 | openstackgerrit | Markus Zoeller (markus_z) proposed openstack/nova master: docs: Explain the flow of the "serial console" feature https://review.openstack.org/476188 | |
| 13:17:19 | mriedem | cdent: so i don't think rebuild_claim will update allocations at all | |
| 13:17:31 | sdague | mriedem: the contract with the user is forced_down means they killed that compute | |
| 13:17:37 | gibi | mriedem: if the compute is stull running but the admin set force-down then it is a user error | |
| 13:17:49 | gibi | admin should fence first then set forced-down flag | |
| 13:17:51 | mriedem | rebuild_call calls _move_call which calls _update_usage_from_migration which calls _update_usage which doesn't call the report client | |
| 13:17:52 | sdague | it is only meant to be used if they've taken that system out of communication | |
| 13:18:15 | sdague | agree with gibi, that's admin error, and we've never attempted to correct for that | |
| 13:18:33 | mriedem | that's not documented anywhere https://developer.openstack.org/api-ref/compute/#update-forced-down | |
| 13:18:47 | sdague | mriedem: ok, well we should document it, that was the whole point of that feature | |
| 13:19:05 | sdague | for HA systems to override nova when it knew better | |
| 13:19:12 | cdent | mriedem: yeah, it looks like the allocation creation is all happening outside the various _claim* methods | |
| 13:19:15 | mriedem | i had reported a bug related to docs on this at one point https://bugs.launchpad.net/nova/+bug/1691871 | |
| 13:19:16 | openstack | Launchpad bug 1691871 in OpenStack Compute (nova) "forced-down vs service disable is not documented well in the compute API reference" [Medium,Confirmed] | |
| 13:19:16 | gibi | mriedem: it is at least in the original spec https://specs.openstack.org/openstack/nova-specs/specs/liberty/implemented/mark-host-down.html | |
| 13:19:18 | cdent | which is somewhat weird | |
| 13:20:28 | cdent | but probably good given what we want eventually | |
| 13:21:25 | sdague | mriedem: sure but as you know, there aren't many idle folks looking for bugs to fix. Docs bugs mostly languish in our tracker. | |
| 13:21:36 | mriedem | gibi: ok so the doctor project is supposed to get an alarm that something is wrong with the host, fence it and then force it down and start evacuating? | |
| 13:21:49 | mriedem | sdague: i planned on fixing that docs gap myself | |
| 13:21:53 | mriedem | but $time | |
| 13:21:59 | gibi | mriedem: I think so, yes | |
| 13:22:02 | sdague | mriedem: sure, that's fine | |
| 13:22:03 | mriedem | what does 'fencing' mean in this case? | |
| 13:22:06 | bauzas | mriedem: do we have a devstack change for https://review.openstack.org/#/c/491854/ ? | |
| 13:22:16 | cdent | re: $time [t 1zqO] | |
| 13:22:16 | purplerbot | <cdent> When do we start asking if the concept of PTL, as currently constructed, is sustainable? [2017-08-10 13:21:20.802315] [n 1zqO] | |
| 13:22:20 | mriedem | bauzas: no | |
| 13:22:22 | gibi | mriedem: power off, or cut the network | |
| 13:22:32 | gibi | mriedem: mostly power off via IPMI | |
| 13:22:33 | bauzas | mriedem: okay, I'll write it | |
| 13:22:34 | sdague | what gibi said | |
| 13:22:48 | mriedem | ok powering off would be ideal | |
| 13:23:04 | sdague | mriedem: but, it could be lots of things. They could also decide to network fence the node | |
| 13:23:05 | mriedem | but network is also good so the compute couldn't send changes to the report client (to conductor i mean) | |
| 13:23:19 | mriedem | as long as it can't get to the placement api then that's sufficient | |
| 13:23:29 | cdent | that comment from bauzas reminds me of something I read while reading gerrit messages last night: did some test have to be changed so that it was running one of the filters we have no declared no longer default? | |
| 13:23:33 | gibi | mriedem: I'm not 100% sure doctor also starts the evacuation automatically after force_down | |
| 13:23:53 | mriedem | gibi: ok but some project, maybe watcher | |
| 13:23:58 | gibi | mriedem: sure | |
| 13:23:58 | sdague | mriedem: I can take a spin on the api-ref, if you review my other doc patches :) | |
| 13:24:03 | gibi | mriedem: we have our on internally :) | |
| 13:24:10 | bauzas | what's the problem with evacuation ? | |
| 13:24:19 | mriedem | sdague: i have one change i need to get in today and that's https://review.openstack.org/#/c/491012/ | |
| 13:24:43 | mriedem | bauzas: see https://review.openstack.org/#/c/491012/ | |
| 13:25:39 | bauzas | mriedem: ah, this... | |
| 13:29:05 | mriedem | cdent: ok left more comments in there | |
| 13:29:14 | mriedem | i think i'm ok with this now, or as much as i can be, | |
| 13:29:26 | mriedem | cdent: i'll fix the functional test quick and push that up, just that change | |
| 13:29:55 | mriedem | and then i'll probably spend some time this afternoon working on a functional test for evacuate | |
| 13:30:01 | mriedem | with the service version stuff | |
| 13:30:39 | mriedem | bauzas: devstack sets this | |
| 13:30:40 | mriedem | iniset $NOVA_CONF filter_scheduler enabled_filters "RetryFilter,AvailabilityZoneFilter,ComputeFilter,ComputeCapabilitiesFilter,ImagePropertiesFilter,CoreFilter,RamFilter,DiskFilter" | |
| 13:30:52 | bauzas | mriedem: I know | |
| 13:30:58 | mriedem | nvm, that's only for the fake hypervisor | |
| 13:30:59 | bauzas | mriedem: I'm just writing atm the change to remove this | |
| 13:31:11 | mriedem | which was only used in the large tests, which we don't run anymore | |
| 13:31:23 | bauzas | we run Core and Ram by default AFAIK | |
| 13:31:31 | bauzas | not Disk tho | |
| 13:31:53 | mriedem | FILTERS="RetryFilter,AvailabilityZoneFilter,RamFilter,DiskFilter,ComputeFilter,ComputeCapabilitiesFilter,ImagePropertiesFilter,ServerGroupAntiAffinityFilter,ServerGroupAffinityFilter,SameHostFilter,DifferentHostFilter" | |
| 13:32:08 | mriedem | ^ is what we run by default | |
| 13:32:41 | bauzas | so, we need to remove Ram *and* Diks | |
| 13:32:56 | bauzas | I was just remembering we still run 2 legacy filters | |
| 13:32:59 | mriedem | http://git.openstack.org/cgit/openstack-dev/devstack/tree/lib/nova#n104 | |
| 13:33:02 | mriedem | ok | |
| 13:33:16 | bauzas | anyway, patch is on its way | |
| 13:34:07 | mriedem | sdague: oh right, that docs bug was because cfriesen was asking about this, see the last paragraph or two in that bug report, | |
| 13:34:30 | mriedem | someone had forced the service down, upgraded everything else, and then when they tried to set forced_down=False, it failed with ServiceTooOld | |
| 13:34:50 | mriedem | which i think just caused some confusion in how that all works | |
| 13:35:01 | openstackgerrit | Sean Dague proposed openstack/nova master: Update api-guide and api-ref to be clear about forced-down https://review.openstack.org/492533 | |
| 13:35:07 | mriedem | the ServiceTooOld is i'm pretty sure a 500 out of the PUT /os-services API too | |
| 13:35:40 | sdague | mriedem: yeh, I don't know what the recovery path is on that all, but at least we can be clear on what the state transition into it should look like | |
| 13:35:46 | cdent | does anyone have anything specific for me to do, or shall I continue chasing reviews and randomly testing things by hand? | |
| 13:37:22 | edleafe | cdent: sudo make me a sandwich | |
| 13:38:50 | bauzas | sdague: mriedem: https://review.openstack.org/#/c/492537/ | |
| 13:39:04 | bauzas | ^ devstack change removing legacy filters FTW | |
| 13:39:21 | sdague | bauzas: ok, cool | |
| 13:39:28 | sdague | bauzas: is there a reason we override those in the first place? | |
| 13:39:39 | bauzas | sdague: correct me if I'm wrong but we don't need to modify grenade since it uses default config optoions ? | |
| 13:39:41 | sdague | would we just use defaults? | |
| 13:39:44 | mriedem | sdague: because devstack had the default + SameHost + DifferentHost | |
| 13:39:49 | bauzas | what mriedem said | |
| 13:39:57 | mriedem | tempest has tests for SameHost/DifferentHost filters | |
| 13:39:59 | sdague | ah, cool, good reason | |
| 13:40:00 | mriedem | which aren't defaults | |
| 13:40:02 | bauzas | we could tho make += "SameHost" | |
| 13:40:09 | bauzas | since it's a listopt | |
| 13:40:15 | mriedem | in bash? | |
| 13:40:23 | sdague | bauzas: that doesn't work in setting nova.conf | |
| 13:40:23 | bauzas | ah, right | |
| 13:40:35 | bauzas | I usually do this directly :) | |
| 13:41:01 | bauzas | sdague: anyway, my question still remains wrt grenade | |
| 13:41:13 | bauzas | sdague: do we need to s// something when we upgrade the node ? | |
| 13:41:23 | sdague | bauzas: grenade should be fine as long as those things didn't get deleted | |
| 13:41:25 | bauzas | in terms of nova.conf -ism | |
| 13:41:36 | sdague | it will use the ocata config in pike | |
| 13:41:46 | bauzas | yeah, my thoughts | |
| 13:41:48 | bauzas | not super then | |