| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-07-27 | |||
| 18:38:55 | jaypipes | dansmith: so we're guaranteed that the only allocations returned are instances that are "on the source host" according to placement. | |
| 18:39:09 | dansmith | mriedem: well the healing doesn't happen all the time, as he pointed out yesterday, only when instances are added/removed, right? | |
| 18:39:22 | dansmith | mriedem: so I'm not sure we'll actually heal over stuff that ocata computes don't do :/ | |
| 18:39:56 | dansmith | jaypipes: delete_allocation_for_instance() operates only on an instance uuid | |
| 18:40:16 | dansmith | jaypipes: if we call that after the flip happened for whatever reason we'll do the wrong thing | |
| 18:40:31 | dansmith | jaypipes: even if we started ten minutes ago and got blocked for a while or something | |
| 18:40:47 | dansmith | jaypipes: really we should use generation on delete to make sure we don't delete something we're not intending to :/ | |
| 18:40:49 | jaypipes | dansmith: right, but didn't you want me to check to see if the allocation had the source compute node UUID in it and if not, don't call delete allocation? | |
| 18:41:08 | dansmith | jaypipes: yes, but delete allocation is only fetching via instance uuid | |
| 18:41:21 | dansmith | https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L999-L1001 | |
| 18:41:36 | jaypipes | dansmith: I understand that, but get_allocations_for_resource_provider() will always only return allocations that "have the source compute node UUID in them" | |
| 18:41:53 | jaypipes | dansmith: so that logic isn't going to filter anything out... | |
| 18:41:59 | mriedem | hangout? | |
| 18:42:05 | dansmith | jaypipes: yeah dude I get that, but we query that way, then time passes, then we delete | |
| 18:42:50 | jaypipes | mriedem: sure | |
| 18:43:09 | dansmith | https://hangouts.google.com/call/5pmzfm5wpfckpptt4l5hxjw5cyu | |
| 18:45:04 | bauzas | can I join ? | |
| 18:45:09 | bauzas | :) | |
| 18:45:13 | bauzas | need more context | |
| 18:45:20 | mriedem | everyone can join | |
| 18:46:53 | smcginnis | Not if you're in China. | |
| 18:46:57 | smcginnis | :) | |
| 18:47:02 | mriedem | nothing stopping them | |
| 18:47:09 | mriedem | climb that mountain | |
| 18:47:32 | smcginnis | :D | |
| 18:52:17 | mriedem | https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L667 | |
| 18:52:30 | mriedem | https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L1030 | |
| 18:52:41 | mriedem | if instance.vm_state not in vm_states.ALLOW_RESOURCE_REMOVAL: | |
| 18:52:49 | dansmith | https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L1007 | |
| 18:56:51 | mriedem | on unshelve we set the host/node on the instance here https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L4423 | |
| 18:57:01 | mriedem | and change the vm_state here https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L4440 | |
| 19:01:30 | openstackgerrit | Merged openstack/nova master: stabilize test_create_delete_server functional test https://review.openstack.org/487772 | |
| 19:11:50 | mriedem | https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L797 | |
| 19:13:24 | mriedem | https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L1012 | |
| 19:14:45 | mriedem | instance_claim updating allocations https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L224 | |
| 19:16:13 | openstackgerrit | OpenStack Proposal Bot proposed openstack/nova master: Updated from global requirements https://review.openstack.org/488034 | |
| 19:19:10 | openstackgerrit | OpenStack Proposal Bot proposed openstack/os-vif master: Updated from global requirements https://review.openstack.org/488086 | |
| 19:21:33 | openstackgerrit | OpenStack Proposal Bot proposed openstack/python-novaclient master: Updated from global requirements https://review.openstack.org/488125 | |
| 19:22:40 | openstackgerrit | Eric Fried proposed openstack/nova master: nova.utils.get_endpoint_data() https://review.openstack.org/488137 | |
| 19:23:37 | efried | mriedem jaypipes mordred Having started poking at the cinderclient construction, I think ^this^ may be a better alternative to get_service_url | |
| 19:23:57 | mriedem | efried: the house is on fire | |
| 19:24:03 | dansmith | mriedem: jaypipes https://etherpad.openstack.org/p/y9sUcb6XW6 | |
| 19:24:17 | mriedem | efried: have sdague check out the service catalog stuff | |
| 19:24:21 | mriedem | he knows more about that than i do | |
| 19:24:24 | efried | mriedem Roger wilco. | |
| 19:29:07 | mordred | efried: yes - that's a great approach | |
| 19:29:35 | efried | mordred Cool, thanks for looking. | |
| 19:31:29 | mriedem | https://review.openstack.org/#/c/244489/ | |
| 19:37:00 | cfriesen_ | jaypipes: did you ever get anywhere with the issue we discussed at the end of June around duplicate scsi device numbers when using virtio-scsi? | |
| 19:37:28 | jaypipes | cfriesen_: nope :( | |
| 19:37:34 | mriedem | cfriesen_: the house is on fire | |
| 19:37:39 | cfriesen_ | jaypipes: I think bug 1702999 is related, as is the "cannot attach new volume to an instance" thread on the openstack-operators list | |
| 19:37:41 | openstack | bug 1702999 in OpenStack Compute (nova) "Can't attach volume if instance boot from volume and virtio-scsi is enabled in the image" [Undecided,Incomplete] https://launchpad.net/bugs/1702999 | |
| 19:38:12 | jaypipes | oh wait, yeah I think we did have a patch for that... | |
| 19:38:36 | jaypipes | cfriesen_: gimme a while... on call | |
| 19:38:57 | mriedem | cfriesen_: this? https://review.openstack.org/#/q/topic:bug/1686116 | |
| 19:42:12 | cfriesen_ | mriedem: looks like it might help. in the case I looked at it would boot (using sda) but trying to attach volumes would fail. | |
| 19:42:58 | cfriesen_ | might be the case that 1702999 is already fixed | |
| 19:49:10 | cdent | jaypipes: if you end up with something that has lose ends by the time you go to bed, feel free to let me know the state of things and I can poke in my morning | |
| 19:49:31 | jaypipes | cdent: thx Chris, will do. | |
| 19:57:30 | openstackgerrit | OpenStack Proposal Bot proposed openstack/python-novaclient master: Updated from global requirements https://review.openstack.org/488125 | |
| 19:59:18 | openstackgerrit | Doug Hellmann proposed openstack/nova master: add a redirect for the old cells landing page https://review.openstack.org/487932 | |
| 20:10:00 | openstackgerrit | OpenStack Proposal Bot proposed openstack/nova master: Updated from global requirements https://review.openstack.org/488034 | |
| 20:12:24 | jaypipes | dansmith: fuuug... so confirm_resize() doesn't run on the destination host. It runs on the source host. :( | |
| 20:12:43 | mriedem | yeah it doesn't call back into rt | |
| 20:12:49 | mriedem | _prep_resize is on dest host right? | |
| 20:12:53 | mriedem | confirm just cleans up shit on the source | |
| 20:12:59 | mriedem | *i think* | |
| 20:13:13 | jaypipes | mriedem: yeah. and that's not the stage of the move operation that we want to have the destination host call PUT /allocations :( | |
| 20:13:31 | mriedem | and i think revert on the source doesn't do anything either, since it's already got resources claimed | |
| 20:13:35 | mriedem | so there is nothing to unclaim | |
| 20:14:10 | bauzas | folks, I will have to bail out, but I'll look at the IRC channel tomorrow morning | |
| 20:14:40 | mriedem | i'm just about to push a change to add some logging and crap in the scheduler.reportclient.delete_allocations_for_instance to sanity check the allocations before we blow them away, to at least see if we're hitting weird stuff in there during migrate tests | |
| 20:14:43 | mriedem | bauzas: o/ | |
| 20:14:44 | bauzas | fer sur, if you need my review, lemme know | |
| 20:15:07 | dansmith | jaypipes: ugh | |
| 20:15:08 | openstackgerrit | OpenStack Proposal Bot proposed openstack/python-novaclient master: Updated from global requirements https://review.openstack.org/488125 | |
| 20:15:47 | sdague | efried: link me | |
| 20:16:06 | efried | sdague https://review.openstack.org/488137 bam | |
| 20:16:10 | efried | sdague TIA. | |
| 20:16:26 | sdague | mriedem: https://review.openstack.org/#/c/487860/ - nova-manage list_cells enhancement | |
| 20:16:30 | sdague | with working tests | |
| 20:16:36 | jaypipes | dansmith, mriedem: so this means really the only thing we can do during confirm_resize() (since it's on the source host) is recalculate the allocation (which will be the doubled-up thing) on the source host RT and remove all entries in the allocation set that refer to the source compute host UUID | |
| 20:16:36 | sdague | I will keep bugging you about it :) | |
| 20:18:03 | dansmith | jaypipes: yeah | |
| 20:18:16 | dansmith | jaypipes: I was thinking something different, but that's smarter :) | |
| 20:19:39 | sdague | efried: I'm surprised this is 'glance' and not 'image' - https://review.openstack.org/#/c/488137/1/nova/image/glance.py@127 | |
| 20:20:39 | efried | sdague It's the conf group name, which needs to correspond to the project name. | |
| 20:20:46 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Sanity check delete_allocation_for_instance https://review.openstack.org/488187 | |
| 20:20:53 | mriedem | dansmith: jaypipes: cdent: ^ | |
| 20:20:58 | mriedem | just for testing at this point | |
| 20:20:58 | efried | sdague Which we then look up in service-types-authority to get the service_type, which is indeed `image` | |
| 20:21:27 | sdague | efried: ok, I was surprised we couldn't just call it image to start with, but if that's how it is, that's fine | |
| 20:21:36 | jaypipes | mriedem: coo. | |
| 20:21:56 | sdague | efried: do we have a test job running this with api_servers not set in devstack? | |
| 20:22:18 | mriedem | jaypipes: "is recalculate the allocation (which will be the doubled-up thing) on the source host RT and remove all entries in the allocation set that refer to the source compute host UUID" is i thought what dansmith and i were talking about earlier, | |
| 20:22:25 | mriedem | which is similar to what my patch is checkingfor | |
| 20:22:43 | efried | sdague Yeah, now that you're saying it, I admit it feels a tad weird. But the point is that nova.utils.get_endpoint_data needs to be able to use that param to find the appropriate conf to load, as well as to find the service_type if it's not specified in the conf. | |
| 20:22:48 | dansmith | mriedem: well, I was assuming we could do it on the destination host | |
| 20:23:03 | dansmith | mriedem: but it doesn't really matter, so yes it's pretty much what we were saying | |