Earlier  
Posted Nick Remark
#openstack-nova - 2017-07-27
18:37:03 jaypipes dansmith: ack
18:37:07 mriedem dansmith: yeah, that seems simplest too
18:37:16 dansmith anybody else looking forward to most of this code going away? :)
18:37:21 mriedem o/
18:37:28 bauzas so since we update the instance.host once the migration is done, the target RT doesn't see it
18:37:31 mriedem so here is another question, which no one is going to like
18:37:44 mriedem this is going to be a change in how the compute is behaving
18:37:48 mriedem and ocata computes won't have this
18:37:50 mriedem so,
18:38:13 jaypipes dansmith: hmm...
18:38:20 mriedem do we (1) make claims in the scheduler dependent on pike computes, or (2) throw to the wind and rely on the dest periodic self-heal fixing the allocations?
18:38:35 jaypipes dansmith: so we're calling get_allocations_for_resource_)provider() and passing in the source compute node UUID.
18:38:55 jaypipes dansmith: so we're guaranteed that the only allocations returned are instances that are "on the source host" according to placement.
18:39:09 dansmith mriedem: well the healing doesn't happen all the time, as he pointed out yesterday, only when instances are added/removed, right?
18:39:22 dansmith mriedem: so I'm not sure we'll actually heal over stuff that ocata computes don't do :/
18:39:56 dansmith jaypipes: delete_allocation_for_instance() operates only on an instance uuid
18:40:16 dansmith jaypipes: if we call that after the flip happened for whatever reason we'll do the wrong thing
18:40:31 dansmith jaypipes: even if we started ten minutes ago and got blocked for a while or something
18:40:47 dansmith jaypipes: really we should use generation on delete to make sure we don't delete something we're not intending to :/
18:40:49 jaypipes dansmith: right, but didn't you want me to check to see if the allocation had the source compute node UUID in it and if not, don't call delete allocation?
18:41:08 dansmith jaypipes: yes, but delete allocation is only fetching via instance uuid
18:41:21 dansmith https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L999-L1001
18:41:36 jaypipes dansmith: I understand that, but get_allocations_for_resource_provider() will always only return allocations that "have the source compute node UUID in them"
18:41:53 jaypipes dansmith: so that logic isn't going to filter anything out...
18:41:59 mriedem hangout?
18:42:05 dansmith jaypipes: yeah dude I get that, but we query that way, then time passes, then we delete
18:42:50 jaypipes mriedem: sure
18:43:09 dansmith https://hangouts.google.com/call/5pmzfm5wpfckpptt4l5hxjw5cyu
18:45:04 bauzas can I join ?
18:45:09 bauzas :)
18:45:13 bauzas need more context
18:45:20 mriedem everyone can join
18:46:53 smcginnis Not if you're in China.
18:46:57 smcginnis :)
18:47:02 mriedem nothing stopping them
18:47:09 mriedem climb that mountain
18:47:32 smcginnis :D
18:52:17 mriedem https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L667
18:52:30 mriedem https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L1030
18:52:41 mriedem if instance.vm_state not in vm_states.ALLOW_RESOURCE_REMOVAL:
18:52:49 dansmith https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L1007
18:56:51 mriedem on unshelve we set the host/node on the instance here https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L4423
18:57:01 mriedem and change the vm_state here https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L4440
19:01:30 openstackgerrit Merged openstack/nova master: stabilize test_create_delete_server functional test https://review.openstack.org/487772
19:11:50 mriedem https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L797
19:13:24 mriedem https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L1012
19:14:45 mriedem instance_claim updating allocations https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L224
19:16:13 openstackgerrit OpenStack Proposal Bot proposed openstack/nova master: Updated from global requirements https://review.openstack.org/488034
19:19:10 openstackgerrit OpenStack Proposal Bot proposed openstack/os-vif master: Updated from global requirements https://review.openstack.org/488086
19:21:33 openstackgerrit OpenStack Proposal Bot proposed openstack/python-novaclient master: Updated from global requirements https://review.openstack.org/488125
19:22:40 openstackgerrit Eric Fried proposed openstack/nova master: nova.utils.get_endpoint_data() https://review.openstack.org/488137
19:23:37 efried mriedem jaypipes mordred Having started poking at the cinderclient construction, I think ^this^ may be a better alternative to get_service_url
19:23:57 mriedem efried: the house is on fire
19:24:03 dansmith mriedem: jaypipes https://etherpad.openstack.org/p/y9sUcb6XW6
19:24:17 mriedem efried: have sdague check out the service catalog stuff
19:24:21 mriedem he knows more about that than i do
19:24:24 efried mriedem Roger wilco.
19:29:07 mordred efried: yes - that's a great approach
19:29:35 efried mordred Cool, thanks for looking.
19:31:29 mriedem https://review.openstack.org/#/c/244489/
19:37:00 cfriesen_ jaypipes: did you ever get anywhere with the issue we discussed at the end of June around duplicate scsi device numbers when using virtio-scsi?
19:37:28 jaypipes cfriesen_: nope :(
19:37:34 mriedem cfriesen_: the house is on fire
19:37:39 cfriesen_ jaypipes: I think bug 1702999 is related, as is the "cannot attach new volume to an instance" thread on the openstack-operators list
19:37:41 openstack bug 1702999 in OpenStack Compute (nova) "Can't attach volume if instance boot from volume and virtio-scsi is enabled in the image" [Undecided,Incomplete] https://launchpad.net/bugs/1702999
19:38:12 jaypipes oh wait, yeah I think we did have a patch for that...
19:38:36 jaypipes cfriesen_: gimme a while... on call
19:38:57 mriedem cfriesen_: this? https://review.openstack.org/#/q/topic:bug/1686116
19:42:12 cfriesen_ mriedem: looks like it might help. in the case I looked at it would boot (using sda) but trying to attach volumes would fail.
19:42:58 cfriesen_ might be the case that 1702999 is already fixed
19:49:10 cdent jaypipes: if you end up with something that has lose ends by the time you go to bed, feel free to let me know the state of things and I can poke in my morning
19:49:31 jaypipes cdent: thx Chris, will do.
19:57:30 openstackgerrit OpenStack Proposal Bot proposed openstack/python-novaclient master: Updated from global requirements https://review.openstack.org/488125
19:59:18 openstackgerrit Doug Hellmann proposed openstack/nova master: add a redirect for the old cells landing page https://review.openstack.org/487932
20:10:00 openstackgerrit OpenStack Proposal Bot proposed openstack/nova master: Updated from global requirements https://review.openstack.org/488034
20:12:24 jaypipes dansmith: fuuug... so confirm_resize() doesn't run on the destination host. It runs on the source host. :(
20:12:43 mriedem yeah it doesn't call back into rt
20:12:49 mriedem _prep_resize is on dest host right?
20:12:53 mriedem confirm just cleans up shit on the source
20:12:59 mriedem *i think*
20:13:13 jaypipes mriedem: yeah. and that's not the stage of the move operation that we want to have the destination host call PUT /allocations :(
20:13:31 mriedem and i think revert on the source doesn't do anything either, since it's already got resources claimed
20:13:35 mriedem so there is nothing to unclaim
20:14:10 bauzas folks, I will have to bail out, but I'll look at the IRC channel tomorrow morning
20:14:40 mriedem i'm just about to push a change to add some logging and crap in the scheduler.reportclient.delete_allocations_for_instance to sanity check the allocations before we blow them away, to at least see if we're hitting weird stuff in there during migrate tests
20:14:43 mriedem bauzas: o/
20:14:44 bauzas fer sur, if you need my review, lemme know
20:15:07 dansmith jaypipes: ugh
20:15:08 openstackgerrit OpenStack Proposal Bot proposed openstack/python-novaclient master: Updated from global requirements https://review.openstack.org/488125
20:15:47 sdague efried: link me
20:16:06 efried sdague https://review.openstack.org/488137 bam
20:16:10 efried sdague TIA.
20:16:26 sdague mriedem: https://review.openstack.org/#/c/487860/ - nova-manage list_cells enhancement
20:16:30 sdague with working tests
20:16:36 sdague I will keep bugging you about it :)
20:16:36 jaypipes dansmith, mriedem: so this means really the only thing we can do during confirm_resize() (since it's on the source host) is recalculate the allocation (which will be the doubled-up thing) on the source host RT and remove all entries in the allocation set that refer to the source compute host UUID
20:18:03 dansmith jaypipes: yeah
20:18:16 dansmith jaypipes: I was thinking something different, but that's smarter :)
20:19:39 sdague efried: I'm surprised this is 'glance' and not 'image' - https://review.openstack.org/#/c/488137/1/nova/image/glance.py@127
20:20:39 efried sdague It's the conf group name, which needs to correspond to the project name.

Earlier   Later