| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-10 | |||
| 10:04:36 | maciejjozefczyk | cdent: :) thx | |
| 10:42:19 | openstackgerrit | Sean Dague proposed openstack/nova master: Clarify that vlan feature means nova-network support https://review.openstack.org/478551 | |
| 10:54:00 | openstackgerrit | Merged openstack/nova master: Remove ram/disk sched filters from default list https://review.openstack.org/491854 | |
| 10:57:37 | openstackgerrit | Merged openstack/nova master: Mark Chance and Caching schedulers as deprecated https://review.openstack.org/492210 | |
| 11:10:07 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: replace chance with filter scheduler in func tests https://review.openstack.org/491529 | |
| 11:12:57 | openstackgerrit | Merged openstack/python-novaclient master: Remove substitutions for command error msg https://review.openstack.org/490705 | |
| 11:26:39 | openstackgerrit | Merged openstack/nova master: [placement] Avoid error log on 405 response https://review.openstack.org/490021 | |
| 11:26:40 | openstackgerrit | Vladyslav Drok proposed openstack/nova master: [placement] Add api-ref for allocation_candidates https://review.openstack.org/481112 | |
| 11:27:14 | openstackgerrit | Vladyslav Drok proposed openstack/nova master: [placement] Make placement_api_docs.py failing https://review.openstack.org/480924 | |
| 11:27:41 | openstackgerrit | Merged openstack/nova master: api-ref: fix security_groups response parameter in os-security-groups https://review.openstack.org/489274 | |
| 11:28:19 | openstackgerrit | Merged openstack/nova master: api-ref: requested security groups are not applied to pre-existing ports https://review.openstack.org/489275 | |
| 11:34:37 | openstackgerrit | OpenStack Proposal Bot proposed openstack/nova master: Imported Translations from Zanata https://review.openstack.org/477091 | |
| 11:39:04 | openstackgerrit | Merged openstack/nova master: Remove translation of log messages https://review.openstack.org/466637 | |
| 11:39:48 | maciejjozefczyk | sdague: Could you check this one? https://review.openstack.org/#/c/491808/ please? | |
| 11:39:49 | openstackgerrit | Merged openstack/nova master: remove mox from unit/virt/vmwareapi/test_driver_api.py https://review.openstack.org/452128 | |
| 11:41:22 | sdague | maciejjozefczyk: how does this handle the synchronization problem that now computes might be trying to delete the same instances at the same time? | |
| 11:41:39 | sdague | previously, by being host scoped, this was a sharded problem | |
| 11:46:29 | openstackgerrit | Merged openstack/nova master: imagebackend: cleanup constructor args to Rbd https://review.openstack.org/490499 | |
| 11:47:13 | openstackgerrit | Merged openstack/nova master: Add policy granularity to the Flavors API https://review.openstack.org/449288 | |
| 11:49:29 | sdague | ah, I see now | |
| 11:49:47 | maciejjozefczyk | sdague: The problem is about instance (which is deleted from nova side) is still running on compute A, but nova says that its deleted and it was on host B | |
| 11:50:00 | sdague | maciejjozefczyk: yeh, I get the problem | |
| 11:50:19 | sdague | I was just trying to make sure that this didn't make it so that multiple computes were trying to delete the same instance | |
| 11:50:48 | sdague | but _get_instances_on_driver should keep us sharded | |
| 11:51:06 | maciejjozefczyk | yes | |
| 11:52:13 | maciejjozefczyk | I'm going to work on patch to rollback migration if deletion of instance will be triggered, in near future | |
| 11:52:22 | sdague | cool | |
| 11:52:49 | maciejjozefczyk | but this fix for already 'lost' and 'working' zombiee instances i think should be in nova | |
| 11:53:30 | maciejjozefczyk | in my installation I have hundreds of them | |
| 11:59:23 | sdague | maciejjozefczyk: yep, +2 on this fix | |
| 12:00:34 | maciejjozefczyk | sdague: thx | |
| 12:06:29 | openstackgerrit | Alex Xu proposed openstack/nova master: placement: avoid returning duplicated alloc_reqs when no sharing rp https://review.openstack.org/492395 | |
| 12:06:30 | openstackgerrit | Alex Xu proposed openstack/nova master: placement: ensure RP maps to those RPs that share with it https://review.openstack.org/480379 | |
| 12:06:36 | alex_xu | cdent: thanks | |
| 12:06:46 | openstackgerrit | Ilya Popov proposed openstack/nova master: Tests: Add cleanup of 'instances' directory https://review.openstack.org/491589 | |
| 12:08:42 | sdague | alex_xu: can I tempt you with doc patches? :) | |
| 12:08:56 | sdague | mostly I'd like to get the manuals stuff merged before I go on vacation next week | |
| 12:09:30 | sdague | https://review.openstack.org/#/q/status:open+project:openstack/nova+branch:master+topic:bp/doc-migration | |
| 12:10:09 | alex_xu | sdague: yea, let me try | |
| 12:13:03 | sdague | alex_xu: thank you | |
| 12:20:24 | openstackgerrit | Chris Dent proposed openstack/nova master: placement: ensure RP maps to those RPs that share with it https://review.openstack.org/480379 | |
| 12:54:14 | openstackgerrit | Lee Yarwood proposed openstack/nova master: libvirt: Reconnect volumes and encryptors during a hard reboot https://review.openstack.org/400384 | |
| 12:54:28 | openstackgerrit | Lee Yarwood proposed openstack/nova master: compute: Detach volumes on _rebuild_default_impl failure https://review.openstack.org/442105 | |
| 12:59:34 | cdent | mriedem: i was partly thinking in terms of “don’t add more churn to zuul, now” | |
| 13:01:59 | mriedem | mmm zuul churn | |
| 13:02:37 | cdent | fresh and tasty | |
| 13:02:41 | mriedem | artom: you love evacuate right? | |
| 13:02:54 | mriedem | gibi: you love finding bugs right? | |
| 13:05:47 | gibi | mriedem: I would put it I like finding them now than getting it from production :) | |
| 13:05:52 | artom | mriedem, in the same way I love, err... | |
| 13:06:01 | artom | Crap, it's too early for witty wordplay | |
| 13:06:08 | artom | mriedem, anyways, what's up? | |
| 13:06:22 | mriedem | my main worry evacuate from an ocata compute messing this up https://review.openstack.org/#/c/491012/ | |
| 13:06:27 | mriedem | artom: i don't know how much you've followed this | |
| 13:06:37 | mriedem | but basically the filter scheduler creates allocations in placement now, | |
| 13:06:43 | mriedem | on both the source and dest computes during a move | |
| 13:06:45 | mriedem | like evacuate | |
| 13:07:28 | mriedem | the problem is that the resource tracker has no concept of other providers than itself, so during it's periodic accounting updates, it overwrites allocations in placement for any other provider | |
| 13:07:37 | mriedem | that patch ^ attempts to resolve that | |
| 13:08:03 | mriedem | by using a minimum compute service version check - so once all of the computes are pike, it will stop doing it's local accounting | |
| 13:08:10 | mriedem | and overwriting the stuff the scheduler created | |
| 13:08:54 | mriedem | one of my worries is that we have an ocata compute that is forced-down, which takes it out of the service version check, but could still be running and trampling on things | |
| 13:09:47 | mriedem | i think it's probably a small window because if you are forcing a compute down and evacuating from it, (1) you're likely to stop that host at some point and (2) once the instances move, the dest compute should be accounting for them - and the scheduler will also do that | |
| 13:10:43 | artom | I'm fuzzy on the resource tracker having no concept of other providers than itself | |
| 13:11:02 | artom | I thought compute nodes were resource providers? | |
| 13:11:21 | mriedem | they are | |
| 13:11:32 | openstackgerrit | Merged openstack/nova master: Keep the code consistent https://review.openstack.org/490304 | |
| 13:12:04 | mriedem | cdent: am i correct in saying that https://review.openstack.org/#/c/491012/ only applies to the periodic update_available_resource task? | |
| 13:12:15 | mriedem | looks like that's the only place that _update_usage_from_instances is called from | |
| 13:12:19 | openstackgerrit | Merged openstack/nova master: add description about key_name https://review.openstack.org/489525 | |
| 13:13:07 | cdent | mriedem: yes | |
| 13:13:10 | mriedem | just thinking that if that code thinks everything is pike and doesn't auto-heal, | |
| 13:13:23 | mriedem | and it misses the instance moving from the ocata compute, | |
| 13:13:31 | mriedem | then we have to be sure that the rebuild_claim handles it | |
| 13:14:44 | cdent | murgh | |
| 13:14:54 | gibi | mriedem: if a compute host is forced_down it should mean that that compute host is fenced | |
| 13:15:16 | gibi | mriedem: therefore it cannot tramp on allocations | |
| 13:15:50 | mriedem | gibi: that doesn't mean the nova-compute service is not running on that host | |
| 13:16:11 | mriedem | and if the service is running, it's update_available_resource periodic is running and could be overwriting allocations for the instance that's being evacuated | |
| 13:16:40 | mriedem | the forced_down flag doesn't do anything besides let the evacuate API proceed before the servicegroup api checkin says the compute is down | |
| 13:16:55 | openstackgerrit | Markus Zoeller (markus_z) proposed openstack/nova master: docs: Explain the flow of the "serial console" feature https://review.openstack.org/476188 | |
| 13:17:19 | mriedem | cdent: so i don't think rebuild_claim will update allocations at all | |
| 13:17:31 | sdague | mriedem: the contract with the user is forced_down means they killed that compute | |
| 13:17:37 | gibi | mriedem: if the compute is stull running but the admin set force-down then it is a user error | |
| 13:17:49 | gibi | admin should fence first then set forced-down flag | |
| 13:17:51 | mriedem | rebuild_call calls _move_call which calls _update_usage_from_migration which calls _update_usage which doesn't call the report client | |
| 13:17:52 | sdague | it is only meant to be used if they've taken that system out of communication | |
| 13:18:15 | sdague | agree with gibi, that's admin error, and we've never attempted to correct for that | |
| 13:18:33 | mriedem | that's not documented anywhere https://developer.openstack.org/api-ref/compute/#update-forced-down | |
| 13:18:47 | sdague | mriedem: ok, well we should document it, that was the whole point of that feature | |
| 13:19:05 | sdague | for HA systems to override nova when it knew better | |
| 13:19:12 | cdent | mriedem: yeah, it looks like the allocation creation is all happening outside the various _claim* methods | |
| 13:19:15 | mriedem | i had reported a bug related to docs on this at one point https://bugs.launchpad.net/nova/+bug/1691871 | |
| 13:19:16 | openstack | Launchpad bug 1691871 in OpenStack Compute (nova) "forced-down vs service disable is not documented well in the compute API reference" [Medium,Confirmed] | |
| 13:19:16 | gibi | mriedem: it is at least in the original spec https://specs.openstack.org/openstack/nova-specs/specs/liberty/implemented/mark-host-down.html | |
| 13:19:18 | cdent | which is somewhat weird | |
| 13:20:28 | cdent | but probably good given what we want eventually | |
| 13:21:25 | sdague | mriedem: sure but as you know, there aren't many idle folks looking for bugs to fix. Docs bugs mostly languish in our tracker. | |
| 13:21:36 | mriedem | gibi: ok so the doctor project is supposed to get an alarm that something is wrong with the host, fence it and then force it down and start evacuating? | |
| 13:21:49 | mriedem | sdague: i planned on fixing that docs gap myself | |