| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-08-03 | |||
| 18:02:48 | jgwentworth | I didn't realize they were different and thought I had made a mistake with trying to delete the snapshot first because that change failed the ceph job, whereas leaving it alone had both jobs passing | |
| 18:49:14 | openstackgerrit | Merged openstack/nova master: Reno for notification-transformation-rocky https://review.openstack.org/588403 | |
| 18:55:33 | openstackgerrit | Chris Dent proposed openstack/nova master: [placement] ensure_rc_cache only at start of process https://review.openstack.org/584086 | |
| 19:07:48 | hansmoleman | cfriesen: what does the VIM do for a "crashed" instance in order to recover it? | |
| 19:13:08 | hansmoleman | seems this could go upstream in some form https://github.com/starlingx-staging/stx-nova/commit/71acfeae0d1c59fdc77704527d763bd85a276f9a#diff-77f9348ab09642ba46409b6828af4af0R7696 - looks like it cleans up old files when a previously evacuated source host comes back online and the instances are off it now (and were being resized when the source was evacuated?) | |
| 19:16:02 | hansmoleman | i thought we already had something like this upstream https://github.com/starlingx-staging/stx-nova/commit/71acfeae0d1c59fdc77704527d763bd85a276f9a#diff-77f9348ab09642ba46409b6828af4af0R7967 | |
| 19:17:03 | openstackgerrit | Merged openstack/nova master: [placement] Move resource_class_cache into placement hierarchy https://review.openstack.org/584085 | |
| 19:19:53 | cfriesen | hansmoleman: for a crashed instance it just does a "stop/start" sequence I think. More generally for an ERROR instance it will cycle through gradually more extreme options (reboot, rebuild, evacuate, etc.) depending on the node state | |
| 19:20:36 | hansmoleman | ok and looking at the _cleanup_running_orphan_instances periodic, and the upstream ComputeManager.init_host, it looks like we don't have something for that, | |
| 19:20:55 | hansmoleman | b/c on compute startup we'll only work with instances still in the db and only destroy guests from the hypervisor that have been evacuated to another host | |
| 19:21:12 | hansmoleman | so i case in that case, the compute host went down when the user deleted the instance from the db | |
| 19:21:25 | hansmoleman | then the compute comes back up and the instance isn't in the db but it's consuming resources on the hypervisor | |
| 19:23:44 | cfriesen | hansmoleman: I think we also got running orphan instance from things like failures during migration | |
| 19:24:01 | openstack | Launchpad bug 1285000 in OpenStack Compute (nova) pike "instance data resides on destination node when vm is deleted during live-migration" [Medium,Fix released] - Assigned to Maciej Jozefczyk (maciej.jozefczyk) | |
| 19:24:01 | hansmoleman | so https://bugs.launchpad.net/nova/+bug/1285000 | |
| 19:26:04 | cfriesen | that'd be one possibility. you could also get something like what mdbooth commented on where a live migration never runs "post live migration at destination" and the system gets into a weird state | |
| 19:28:56 | cfriesen | the orphan audit dates from havana, when things were not quite as robust as they are now | |
| 19:30:01 | hansmoleman | ack | |
| 19:39:20 | hansmoleman | cfriesen: interesting https://github.com/starlingx-staging/stx-nova/commit/71acfeae0d1c59fdc77704527d763bd85a276f9a#diff-afb9c0c0ca5276c7eacd987bbf51d8e6R447 | |
| 19:39:48 | hansmoleman | does upstream retrieve the volume_image_metadata properly for volume-backed scheduling with things like the NUMATopologyFilter? | |
| 19:41:53 | hansmoleman | looks like we should, compute API _get_bdm_image_metadata | |
| 19:41:59 | hansmoleman | gets the volume image metadata from the volume | |
| 19:53:02 | openstack | Launchpad bug 1785318 in OpenStack Compute (nova) "evacuate rebuild claim will not use any image_meta for volume-backed instances" [Medium,Triaged] | |
| 19:53:02 | hansmoleman | https://bugs.launchpad.net/nova/+bug/1785318 | |
| 20:05:37 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Remove old check_attach version check in API https://review.openstack.org/588348 | |
| 20:09:43 | hansmoleman | fried_rice: ^ updated the commit message on that for the del instance.id thing in the tests | |
| 20:09:53 | hansmoleman | also updated some really really old comment | |
| 20:14:31 | cfriesen | hansmoleman: thanks for opening that bug report. I believe we have evacuate working reliably with NumaTopology. | |
| 20:14:42 | hansmoleman | yeah i don't see why it wouldn't | |
| 20:14:49 | hansmoleman | i was thinking of the isolated hosts filter, | |
| 20:14:58 | hansmoleman | which relies on the image id which we don't have in the request spec for bfv | |
| 20:15:33 | hansmoleman | https://review.openstack.org/#/c/543263/ | |
| 20:15:37 | cfriesen | hansmoleman: there will likely be other cases where we fixed something and then it's been fixed upstream and we ported the change without retesting against upstream first. | |
| 20:15:54 | hansmoleman | ^ still not sure *why* or if it was intentional that we don't store the RequestSpec.image.id for bfv instances | |
| 20:16:22 | hansmoleman | we also just don't really test the isolated hosts filter | |
| 20:48:21 | hansmoleman | cfriesen: i'll have a patch up for that shortly, slightly different than what's in starlingx | |
| 20:59:00 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Fix image-defined numa claims during evacuate https://review.openstack.org/588657 | |
| 21:22:25 | openstackgerrit | Eric Fried proposed openstack/nova master: Remove redundant _update()s https://review.openstack.org/588091 | |
| 21:29:37 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Optimize AZ lookup during schedule_and_build_instances https://review.openstack.org/588665 | |
| 21:37:05 | hansmoleman | dansmith: looks like another place for the long rpc timeout https://github.com/starlingx-staging/stx-nova/commit/71acfeae0d1c59fdc77704527d763bd85a276f9a#diff-2a50b2dbeb123b515ebb4b917ae1cb2bR751 | |
| 21:47:37 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Use CONF.long_rpc_timeout in post_live_migration_at_destination https://review.openstack.org/588668 | |
| 21:50:53 | hansmoleman | cfriesen: i thought the api change to return the server group that each server is in was kind of interesting, | |
| 21:51:05 | hansmoleman | but that kind of sucks for performance if you're listing servers with details for 1000 servers | |
| 21:51:11 | hansmoleman | since it's an api db query per instance | |
| 21:51:32 | hansmoleman | alternatively that could be done by storing the instance group info in instance_extra with the instance, | |
| 21:51:39 | hansmoleman | or adding a member filter to GET /os-server-groups | |
| 21:51:52 | hansmoleman | so you could just get server groups by a given member server | |
| 21:52:14 | hansmoleman | that list would always return at most 1 since a server can't be in more than one group | |
| 21:55:35 | cfriesen | hansmoleman: that implementation was driven partly by trying to make it as easy to port as possible. | |
| 21:56:12 | hansmoleman | yeah i get that | |
| 21:56:22 | hansmoleman | i think the server group is only returned if the wrs-header is present too... | |
| 21:56:26 | cfriesen | yes] | |
| 21:57:24 | hansmoleman | was also thinking we could do like a GET /servers/{id}/group but that is pretty boring, it would just return the server group subresource | |
| 21:57:46 | hansmoleman | anywho | |
| 21:59:08 | cfriesen | heading out? | |
| 21:59:10 | openstack | Launchpad bug 1781286 in OpenStack Compute (nova) "CantStartEngineError in cell conductor during reschedule - get_host_availability_zone up-call" [Medium,Triaged] | |
| 21:59:10 | hansmoleman | jgwentworth: on https://bugs.launchpad.net/nova/+bug/1781286 i just remembered, | |
| 21:59:15 | hansmoleman | cfriesen: of course not | |
| 21:59:38 | hansmoleman | jgwentworth: one way i thought about fixing that was adding a different migrate_server (or whatever it's called) method in conductor that's not trying to target a cell, | |
| 21:59:59 | hansmoleman | because that's how build_resources works, it's not using the @targets_cell decorator because it's called from the api for scheduling and from the compute for reschedules | |
| 22:00:10 | hansmoleman | so really we should do similar for rescheduling a resize/cold migrate | |
| 22:00:22 | cfriesen | hansmoleman: as a heads-up, I'm on vacation next week. you'll have to save up your questions. :) | |
| 22:00:31 | hansmoleman | cfriesen: damn | |
| 22:00:40 | hansmoleman | i'm leaving for china on friday | |
| 22:00:48 | cfriesen | you can try asking Dean | |
| 22:01:19 | hansmoleman | given i'm gotten through the api extension and nova/compute/* changes so far, i think i'm pretty good | |
| 22:01:30 | hansmoleman | i expect some scheduler stuff but mostly linked to what i've already seen | |
| 22:01:37 | hansmoleman | *i've | |
| 22:01:55 | cfriesen | there are some changes around the server group affinity to close some races in the scheduler | |
| 22:02:49 | hansmoleman | jgwentworth: reason i bring that up now, is if we fixed the bug that way, it's a new rpc api method which isn't backportable, | |
| 22:02:57 | hansmoleman | so we'd have to get that done before rc if we want it in rocky | |
| 22:04:49 | hansmoleman | i also thought about putting the az per host on the Selection object that we pass down for alternates, | |
| 22:04:58 | hansmoleman | but that's an rcp api version bump on the Selection object, so again, not backportable | |
| 22:06:36 | hansmoleman | ew it's actually 2 separate but very similar bugs and we'd need both | |
| 22:06:52 | hansmoleman | @targets_cell looks up the instance mapping which fails if the cell conductor isn't configured with the api db | |
| 22:07:03 | hansmoleman | looking up the host az fails if not configured for the api db | |
| 22:14:01 | hansmoleman | mnaser: i'm assuming you have [api_database]/connection configured in your cell conductors? or are you just using a single global conductor? | |
| 22:15:25 | mnaser | s/has/had/ | |
| 22:15:44 | mnaser | hansmoleman: right now just a global conductor but we'll be switching to cell conductor soon | |
| 22:33:55 | hansmoleman | jgwentworth: just added https://blueprints.launchpad.net/nova/+spec/fix-reschedule-up-calls for stein and to the ptg etherpad; i think it's too much change at this point to try and rush a fix for rocky | |
| 22:34:04 | hansmoleman | especially given it's been broken since pike | |
| 22:36:34 | hansmoleman | cfriesen: hmm looks like a bug upstream https://github.com/starlingx-staging/stx-nova/commit/71acfeae0d1c59fdc77704527d763bd85a276f9a#diff-378a96ec6159d0a2f8ec7ab71bc3843bR921 - i know we change the request spec on resize, but do we revert the request_spec.flavor on revert resize? | |
| 22:37:56 | hansmoleman | another good reason why we can't trust the request spec a lot of the time.. | |
| 22:41:10 | openstackgerrit | Merged openstack/os-vif master: Support for OVS DB TCP socket communication. https://review.openstack.org/587378 | |
| 22:45:21 | openstack | Launchpad bug 1785339 in OpenStack Compute (nova) "RequestSpec.flavor is not reverted on resize revert" [Medium,Triaged] | |
| 22:45:21 | hansmoleman | https://bugs.launchpad.net/nova/+bug/1785339 | |
| 22:45:22 | hansmoleman | busted since newton | |
| 23:14:23 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Update RequestSpec.flavor on resize_revert https://review.openstack.org/588689 | |
| 23:18:05 | openstackgerrit | Merged openstack/nova master: Remove unused flavor_delete_info() method https://review.openstack.org/588621 | |
| 23:24:43 | hansmoleman | cfriesen: i see what you mean about affinity races https://github.com/starlingx-staging/stx-nova/commit/71acfeae0d1c59fdc77704527d763bd85a276f9a#diff-d29c9372baf108a281712642550918dcR88 | |
| 23:24:50 | hansmoleman | should probably report a bug for that | |
| 23:25:18 | hansmoleman | since we don't do the late anti-affinity check on compute like we do for server create and evacuate (which would be up-calls for cells v2 now too) | |
| 23:25:45 | openstack | Launchpad bug 1600251 in OpenStack Compute (nova) "live migration does not honor server group policy" [High,Confirmed] | |
| 23:25:45 | hansmoleman | oh heh https://bugs.launchpad.net/nova/+bug/1600251 | |
| 23:30:59 | hansmoleman | i think part of that is fixed with https://review.openstack.org/#/c/527799/ | |
| 23:31:08 | hansmoleman | but not sure where we re-calculate the group members prior to scheduling | |
| #openstack-nova - 2018-08-04 | |||
| 01:23:51 | openstackgerrit | Merged openstack/nova master: [placement] ensure_rc_cache only at start of process https://review.openstack.org/584086 | |
| 08:07:50 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Refactor NeutronFixture https://review.openstack.org/588338 | |
| 10:37:59 | openstackgerrit | Tetsuro Nakamura proposed openstack/nova master: Adds a test for getting allocations API https://review.openstack.org/588886 | |
| 10:38:00 | openstackgerrit | Tetsuro Nakamura proposed openstack/nova master: Not use project table for user table https://review.openstack.org/588887 | |