Earlier  
Posted Nick Remark
#openstack-nova - 2018-05-03
18:56:49 mriedem looks like some rax/xen thing
18:56:54 mriedem that likely never made the api change upstream
18:56:57 mriedem gd rax
18:57:25 melwitt guh, weird. so we have patches that say "now you can reboot a rescued instance" but you can't because the API kicks you out
18:57:38 mriedem the rax api code probably lets you
18:57:41 melwitt right
18:57:51 mriedem where was that one guy that still works at rax?
18:58:01 mriedem the guy always in korea..
18:58:08 melwitt tbh, I don't know whether we're suppose to be able to reboot rescued instances. I'm not that familiar with the "rescue" function
18:58:18 mriedem neither am i,
18:58:23 mriedem but i can read code :)
18:58:25 mriedem that's how i found this
18:58:52 melwitt well, yeah. I mean whether or not it makes sense. I would guess from the lack of user complaints that it's not usual to try to reboot one
18:59:06 mriedem mikal: johnthetubaguy: when you're around, maybe you can sort this out - did rax have a proprietary change to allow rebooting rescued instances? https://review.openstack.org/#/q/topic:bug/1170237+(status:open+OR+status:merged)
18:59:24 mriedem because the upstream api doesn't allow that
18:59:36 mriedem see https://github.com/openstack/nova/blob/master/nova/compute/vm_states.py#L69-L72
19:09:22 openstackgerrit Matt Riedemann proposed openstack/nova master: Fix being able to hard reboot a pausing instance https://review.openstack.org/566143
19:26:58 openstackgerrit Oliver Walsh proposed openstack/nova master: Fix handling of connect issues in _ensure_resource_provider https://review.openstack.org/566148
19:28:08 mriedem owalsh: i already have a patch up for that
19:28:13 mriedem beat you by a few hours
19:28:27 mriedem https://review.openstack.org/#/c/566096/
19:29:18 owalsh mriedem: snap
19:32:27 owalsh mriedem: worse on pike, it cache None
19:32:42 mriedem yeah it was a backport regression
19:33:58 melwitt oh, so _that's_ where the None RP comes from. geesh
19:34:22 melwitt I remember we reverted a thing while we were trying to get backports done for the stable releases
19:34:48 mriedem yeah that was ocata
19:35:00 mriedem i suspect the fleetify devstack stuff in pike+ was hiding it for us
19:35:01 mriedem somehow
19:35:28 mriedem although even in ocata devstack i thought we started the compute last
19:44:00 mriedem dansmith: what do you think a safe batch size is for this heal allocations CLI? at first i default to CONF.api.max_limit but that's 1000 which seems way too big https://review.openstack.org/#/c/565886/5/nova/cmd/manage.py@1776 - map_instances defaults to 50
19:44:05 mriedem so was thinking about using 50
19:44:21 dansmith yeah 1000 is too much
19:44:24 dansmith 50 is probably good
19:44:25 mriedem online_data_migrations also does 50
19:44:32 dansmith yup
19:44:54 melwitt did we do an audit of other uses of safe_connect where None can be returned?
19:45:11 mriedem melwitt: jaypipes has been working on untangling that
19:45:18 melwitt k, cool
19:59:36 openstackgerrit Takashi NATSUME proposed openstack/nova master: Fix wrong arguments for 'detach_volume' https://review.openstack.org/566152
20:06:49 openstackgerrit Matt Riedemann proposed openstack/nova master: Add nova-manage placement heal_allocations CLI https://review.openstack.org/565886
20:11:38 mriedem owalsh: if you want to do the backport to queens and pike for https://review.openstack.org/#/c/566096/ that would speed things along so i can +2 the backports
20:11:52 mriedem owalsh: beware: the provider tree stuff in rocky will likely mean merge conflicts for the backports
20:16:59 arvindn05 dansmith: looks like there was agreement on the rebuild instance with traits thread. Can you send out an update on ML on the final approach?
20:17:03 openstackgerrit Takashi NATSUME proposed openstack/nova master: Remove mox in virt/test_block_device.py https://review.openstack.org/566153
20:17:43 dansmith arvindn05: I started it a bit ago but got distracted.. if there is agreement you're not blocked right?
20:23:15 openstackgerrit Matt Riedemann proposed openstack/nova master: Create volume attachment during boot from volume in compute https://review.openstack.org/541420
20:24:17 arvindn05 dansmith: i am blocked at this point because my approved patch https://review.openstack.org/#/c/560596/ is also holding based on the decision on the rebuild issue
20:25:15 arvindn05 my next patches would be update the spec with decision on rebuild and propose the code patch for the same
20:25:24 dansmith arvindn05: you mean you're blocked because your patches don't do what mriedem wants yeah?
20:25:48 dansmith arvindn05: if you'd do what we said in the meeting this morning, then he'd remove his -W and everything would move along, AFAICT
20:26:53 arvindn05 dansmith: yup. but i was unfortunately not in the meeting and not entirely sure what approach was decided
20:27:12 dansmith if only there was a log...
20:27:34 arvindn05 dansmith: is it fair to summarize it as we want to go check the allocations route
20:27:50 arvindn05 (8:25:17 AM) efried: arvindn05: That was the impression I got. But yeah, let's see what dansmith has to say.
20:27:59 dansmith arvindn05: yes
20:28:39 arvindn05 dansmith: great. Thanks for confirming....i will start with the spec and the code patch. glad the deadlock was resolved :)
20:28:44 efried ++
20:30:02 mriedem arvindn05: fyi http://eavesdrop.openstack.org/meetings/nova/2018/nova.2018-05-03-14.00.log.html#l-161
20:31:05 openstackgerrit Matt Riedemann proposed openstack/nova master: Fix being able to hard reboot a pausing instance https://review.openstack.org/566143
20:31:57 owalsh mriedem: sure
20:32:20 mriedem thanks, hopefully it's not too bad, the patch is pretty isolated
20:32:40 arvindn05 mriedem: thanks got that from gibi as well :) last statement was <dansmith> I shall commentificate upon the threadage and reviewage so wanted to confirm :)
20:34:57 openstackgerrit Matt Riedemann proposed openstack/nova master: Cleanup placement policy generator docs https://review.openstack.org/565225
20:41:26 arvindn05 dansmith: one other clarification, we want to validate the allocations vs traits always correct? Even when we are rebuilding with the same image and no traits have changed
20:48:55 melwitt I would think so -- while the image hasn't changed, the underlying resource provider traits could have
20:50:12 mriedem arvindn05: that gets tricky depending on where you do the validation,
20:50:31 mriedem if the image doesn't change, i'm not sure if you get to the point in conductor where we'd be calling the scheduler
20:50:47 mriedem i realze you're not calling the scheduler to do the validation, but i assumed it would be in the same block
20:51:03 arvindn05 mriedem: i am thinking it would be in the conductor
20:51:32 arvindn05 conductor already has the placement client and all its dependencies need to make the validation
20:52:13 arvindn05 conductor.manager.ComputeTaskManager#rebuild_instance somewhere in here is where the validation logic would lie
20:55:06 mriedem arvindn05: you'd have an else block here https://github.com/openstack/nova/blob/0ef3c685b9d2e0049f38fcf1a268870e69a5b9cf/nova/conductor/manager.py#L944 when recreate is False
20:55:07 arvindn05 in case of rebuild we would add logic here https://github.com/openstack/nova/blob/master/nova/conductor/manager.py#L897
20:56:26 arvindn05 mriedem: sorry...looking at your code pointer now
20:56:59 mriedem melwitt: so i'm looking at https://review.openstack.org/#/c/540258/ and something is bugging me,
20:57:13 mriedem with an affinity group policy, the group members would all have to be in the same cell because they have to be on the same host
20:57:25 mriedem also soft-affinity might screw with that...
20:57:42 mriedem for an anti-affinity group, we could have instances in different cells i'd thinkg
20:57:43 mriedem *think
20:58:09 melwitt yeah, I think so
20:58:35 melwitt are you suggesting that if policy == affinity then limit to same cell?
20:58:37 mriedem so if i'm doing a move operation on an instance in an anti-affinity group, we have a targeted context for the cell that instance being moved is in,
20:58:50 arvindn05 mriedem: from the comment in line 916 that whole block executes only in case we are doing a rebuild with a new image....
20:59:08 mriedem and we'll only get group hosts in that cell to make sure the instance doesn't go to a host that another member of the same group in the same cell is in, since we can't move across cells
20:59:17 melwitt right
20:59:19 mriedem which is correct
20:59:19 arvindn05 if we want to run the validation always...we would have to do outside the block, correct?
21:00:13 mriedem arvindn05: well, rebuild + new image or evacuate
21:00:18 mriedem evacuate is a rebuild on another new host
21:00:20 mriedem but same image
21:00:35 mriedem in that case, the scheduler will go through GET /allocation_candidates and do a claim,
21:00:47 mriedem so we don't need to do any image validation for traits in conductor for evacuate b/c that would be redundant
21:01:39 arvindn05 yup...i got that for evacuate...how about the case of rebuild without the image being changed
21:01:59 mriedem melwitt: ok and if we're not targeted (not a move, just instance create), then we need to get group hosts from all cells because we don't know in which cell the instance is going to land in
21:02:10 melwitt mriedem: right
21:02:52 mriedem and this happens before we ever run the affinity/anti-affinity filters right?
21:03:26 mriedem yeah called from schedule_and_build_instances
21:03:33 melwitt yes, this is the setup actually pretty early on in compute/api, I think
21:03:51 melwitt okay maybe not (sorry I forgot)

Earlier   Later