| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-10 | |||
| 11:50:48 | sdague | but _get_instances_on_driver should keep us sharded | |
| 11:51:06 | maciejjozefczyk | yes | |
| 11:52:13 | maciejjozefczyk | I'm going to work on patch to rollback migration if deletion of instance will be triggered, in near future | |
| 11:52:22 | sdague | cool | |
| 11:52:49 | maciejjozefczyk | but this fix for already 'lost' and 'working' zombiee instances i think should be in nova | |
| 11:53:30 | maciejjozefczyk | in my installation I have hundreds of them | |
| 11:59:23 | sdague | maciejjozefczyk: yep, +2 on this fix | |
| 12:00:34 | maciejjozefczyk | sdague: thx | |
| 12:06:29 | openstackgerrit | Alex Xu proposed openstack/nova master: placement: avoid returning duplicated alloc_reqs when no sharing rp https://review.openstack.org/492395 | |
| 12:06:30 | openstackgerrit | Alex Xu proposed openstack/nova master: placement: ensure RP maps to those RPs that share with it https://review.openstack.org/480379 | |
| 12:06:36 | alex_xu | cdent: thanks | |
| 12:06:46 | openstackgerrit | Ilya Popov proposed openstack/nova master: Tests: Add cleanup of 'instances' directory https://review.openstack.org/491589 | |
| 12:08:42 | sdague | alex_xu: can I tempt you with doc patches? :) | |
| 12:08:56 | sdague | mostly I'd like to get the manuals stuff merged before I go on vacation next week | |
| 12:09:30 | sdague | https://review.openstack.org/#/q/status:open+project:openstack/nova+branch:master+topic:bp/doc-migration | |
| 12:10:09 | alex_xu | sdague: yea, let me try | |
| 12:13:03 | sdague | alex_xu: thank you | |
| 12:20:24 | openstackgerrit | Chris Dent proposed openstack/nova master: placement: ensure RP maps to those RPs that share with it https://review.openstack.org/480379 | |
| 12:54:14 | openstackgerrit | Lee Yarwood proposed openstack/nova master: libvirt: Reconnect volumes and encryptors during a hard reboot https://review.openstack.org/400384 | |
| 12:54:28 | openstackgerrit | Lee Yarwood proposed openstack/nova master: compute: Detach volumes on _rebuild_default_impl failure https://review.openstack.org/442105 | |
| 12:59:34 | cdent | mriedem: i was partly thinking in terms of “don’t add more churn to zuul, now” | |
| 13:01:59 | mriedem | mmm zuul churn | |
| 13:02:37 | cdent | fresh and tasty | |
| 13:02:41 | mriedem | artom: you love evacuate right? | |
| 13:02:54 | mriedem | gibi: you love finding bugs right? | |
| 13:05:47 | gibi | mriedem: I would put it I like finding them now than getting it from production :) | |
| 13:05:52 | artom | mriedem, in the same way I love, err... | |
| 13:06:01 | artom | Crap, it's too early for witty wordplay | |
| 13:06:08 | artom | mriedem, anyways, what's up? | |
| 13:06:22 | mriedem | my main worry evacuate from an ocata compute messing this up https://review.openstack.org/#/c/491012/ | |
| 13:06:27 | mriedem | artom: i don't know how much you've followed this | |
| 13:06:37 | mriedem | but basically the filter scheduler creates allocations in placement now, | |
| 13:06:43 | mriedem | on both the source and dest computes during a move | |
| 13:06:45 | mriedem | like evacuate | |
| 13:07:28 | mriedem | the problem is that the resource tracker has no concept of other providers than itself, so during it's periodic accounting updates, it overwrites allocations in placement for any other provider | |
| 13:07:37 | mriedem | that patch ^ attempts to resolve that | |
| 13:08:03 | mriedem | by using a minimum compute service version check - so once all of the computes are pike, it will stop doing it's local accounting | |
| 13:08:10 | mriedem | and overwriting the stuff the scheduler created | |
| 13:08:54 | mriedem | one of my worries is that we have an ocata compute that is forced-down, which takes it out of the service version check, but could still be running and trampling on things | |
| 13:09:47 | mriedem | i think it's probably a small window because if you are forcing a compute down and evacuating from it, (1) you're likely to stop that host at some point and (2) once the instances move, the dest compute should be accounting for them - and the scheduler will also do that | |
| 13:10:43 | artom | I'm fuzzy on the resource tracker having no concept of other providers than itself | |
| 13:11:02 | artom | I thought compute nodes were resource providers? | |
| 13:11:21 | mriedem | they are | |
| 13:11:32 | openstackgerrit | Merged openstack/nova master: Keep the code consistent https://review.openstack.org/490304 | |
| 13:12:04 | mriedem | cdent: am i correct in saying that https://review.openstack.org/#/c/491012/ only applies to the periodic update_available_resource task? | |
| 13:12:15 | mriedem | looks like that's the only place that _update_usage_from_instances is called from | |
| 13:12:19 | openstackgerrit | Merged openstack/nova master: add description about key_name https://review.openstack.org/489525 | |
| 13:13:07 | cdent | mriedem: yes | |
| 13:13:10 | mriedem | just thinking that if that code thinks everything is pike and doesn't auto-heal, | |
| 13:13:23 | mriedem | and it misses the instance moving from the ocata compute, | |
| 13:13:31 | mriedem | then we have to be sure that the rebuild_claim handles it | |
| 13:14:44 | cdent | murgh | |
| 13:14:54 | gibi | mriedem: if a compute host is forced_down it should mean that that compute host is fenced | |
| 13:15:16 | gibi | mriedem: therefore it cannot tramp on allocations | |
| 13:15:50 | mriedem | gibi: that doesn't mean the nova-compute service is not running on that host | |
| 13:16:11 | mriedem | and if the service is running, it's update_available_resource periodic is running and could be overwriting allocations for the instance that's being evacuated | |
| 13:16:40 | mriedem | the forced_down flag doesn't do anything besides let the evacuate API proceed before the servicegroup api checkin says the compute is down | |
| 13:16:55 | openstackgerrit | Markus Zoeller (markus_z) proposed openstack/nova master: docs: Explain the flow of the "serial console" feature https://review.openstack.org/476188 | |
| 13:17:19 | mriedem | cdent: so i don't think rebuild_claim will update allocations at all | |
| 13:17:31 | sdague | mriedem: the contract with the user is forced_down means they killed that compute | |
| 13:17:37 | gibi | mriedem: if the compute is stull running but the admin set force-down then it is a user error | |
| 13:17:49 | gibi | admin should fence first then set forced-down flag | |
| 13:17:51 | mriedem | rebuild_call calls _move_call which calls _update_usage_from_migration which calls _update_usage which doesn't call the report client | |
| 13:17:52 | sdague | it is only meant to be used if they've taken that system out of communication | |
| 13:18:15 | sdague | agree with gibi, that's admin error, and we've never attempted to correct for that | |
| 13:18:33 | mriedem | that's not documented anywhere https://developer.openstack.org/api-ref/compute/#update-forced-down | |
| 13:18:47 | sdague | mriedem: ok, well we should document it, that was the whole point of that feature | |
| 13:19:05 | sdague | for HA systems to override nova when it knew better | |
| 13:19:12 | cdent | mriedem: yeah, it looks like the allocation creation is all happening outside the various _claim* methods | |
| 13:19:15 | mriedem | i had reported a bug related to docs on this at one point https://bugs.launchpad.net/nova/+bug/1691871 | |
| 13:19:16 | openstack | Launchpad bug 1691871 in OpenStack Compute (nova) "forced-down vs service disable is not documented well in the compute API reference" [Medium,Confirmed] | |
| 13:19:16 | gibi | mriedem: it is at least in the original spec https://specs.openstack.org/openstack/nova-specs/specs/liberty/implemented/mark-host-down.html | |
| 13:19:18 | cdent | which is somewhat weird | |
| 13:20:28 | cdent | but probably good given what we want eventually | |
| 13:21:25 | sdague | mriedem: sure but as you know, there aren't many idle folks looking for bugs to fix. Docs bugs mostly languish in our tracker. | |
| 13:21:36 | mriedem | gibi: ok so the doctor project is supposed to get an alarm that something is wrong with the host, fence it and then force it down and start evacuating? | |
| 13:21:49 | mriedem | sdague: i planned on fixing that docs gap myself | |
| 13:21:53 | mriedem | but $time | |
| 13:21:59 | gibi | mriedem: I think so, yes | |
| 13:22:02 | sdague | mriedem: sure, that's fine | |
| 13:22:03 | mriedem | what does 'fencing' mean in this case? | |
| 13:22:06 | bauzas | mriedem: do we have a devstack change for https://review.openstack.org/#/c/491854/ ? | |
| 13:22:16 | cdent | re: $time [t 1zqO] | |
| 13:22:16 | purplerbot | <cdent> When do we start asking if the concept of PTL, as currently constructed, is sustainable? [2017-08-10 13:21:20.802315] [n 1zqO] | |
| 13:22:20 | mriedem | bauzas: no | |
| 13:22:22 | gibi | mriedem: power off, or cut the network | |
| 13:22:32 | gibi | mriedem: mostly power off via IPMI | |
| 13:22:33 | bauzas | mriedem: okay, I'll write it | |
| 13:22:34 | sdague | what gibi said | |
| 13:22:48 | mriedem | ok powering off would be ideal | |
| 13:23:04 | sdague | mriedem: but, it could be lots of things. They could also decide to network fence the node | |
| 13:23:05 | mriedem | but network is also good so the compute couldn't send changes to the report client (to conductor i mean) | |
| 13:23:19 | mriedem | as long as it can't get to the placement api then that's sufficient | |
| 13:23:29 | cdent | that comment from bauzas reminds me of something I read while reading gerrit messages last night: did some test have to be changed so that it was running one of the filters we have no declared no longer default? | |
| 13:23:33 | gibi | mriedem: I'm not 100% sure doctor also starts the evacuation automatically after force_down | |
| 13:23:53 | mriedem | gibi: ok but some project, maybe watcher | |
| 13:23:58 | gibi | mriedem: sure | |
| 13:23:58 | sdague | mriedem: I can take a spin on the api-ref, if you review my other doc patches :) | |
| 13:24:03 | gibi | mriedem: we have our on internally :) | |
| 13:24:10 | bauzas | what's the problem with evacuation ? | |