Earlier  
Posted Nick Remark
#openstack-nova - 2017-11-01
14:46:47 efried alex_xu Still around?
14:49:11 openstackgerrit Matt Riedemann proposed openstack/nova master: Raise specific exception when swapping migration allocations fails https://review.openstack.org/517004
14:49:25 dansmith mriedem: hmm, you think there's a bug in that delete-if-deleted logic
14:49:38 mriedem dansmith: yes, that's what's deleting the allocation
14:49:42 mriedem http://logs.openstack.org/08/516708/3/gate/legacy-tempest-dsvm-py35/4d8d6a3/logs/screen-n-cpu.txt.gz#_Oct_31_23_18_00_780729
14:49:57 mriedem https://review.openstack.org/517004 fixes the misleading 400 out of the API
14:51:29 mriedem dansmith: so this must be a race between the time the RT tracks the instance as a 'known instance' and the time the update_available_resource periodic runs
14:52:15 dansmith mriedem: it's not supposed to delete it unless it's really deleted=yes though
14:52:16 mriedem yeah the periodic starts here http://logs.openstack.org/08/516708/3/gate/legacy-tempest-dsvm-py35/4d8d6a3/logs/screen-n-cpu.txt.gz#_Oct_31_23_18_00_165850
14:52:41 dansmith mriedem: although, I wonder if it might create the allocation before it's created in the cell db and thus it's a BR only, compute thinks it's deleted like archived
14:52:56 mriedem eesh
14:54:02 mriedem yeah i guess (1) scheduler creates allocation, (2) periodic on compute starts - gets allocations against itself, deletes allocations b/c InstanceNotFound in cell (3) superconductor creates instance in cell for the selected host
14:54:14 mriedem shite
14:54:43 dansmith yeah
14:54:44 mriedem and...we can't really have the compute try to find out if a build request exists can we given that's API DB and the compute shouldn't have access to the API DB
14:54:55 dansmith nope
14:55:01 mriedem i mean, i guess if you're ocata or pike single-cell that might be something you're still doing
14:55:04 mriedem for the late affinity check
14:55:26 mriedem well i'll open a bug to track it anyway
14:55:28 dansmith we've never had computes use objects that are only on the api db, AFAIK,
14:55:41 mriedem server group is in the api db
14:55:45 mriedem *InstanceGroup
14:55:53 dansmith well, because it moved
14:56:08 dansmith anyway, that's a door we don't want to open, even under config I think
14:56:08 mriedem yeah
14:56:32 dansmith deleting only if we find an actually-deleted instance would be one way to handle it,
14:56:43 dansmith which pushes the race to the delete-archive boundary
14:56:48 dansmith but that's much better, IMHO
14:57:17 dansmith if they're doing fast purging (like immediate purging of all deleted things, then they could bump into it, but much less likely, and we could have nova-manage check for allocations potentially
15:01:54 openstackgerrit Matthew Booth proposed openstack/nova master: libvirt: Don't VIR_MIGRATE_NON_SHARED_INC without migrate_disks https://review.openstack.org/507202
15:02:20 openstackgerrit Matthew Booth proposed openstack/nova master: DNM: Run test_volume_backed_live_migration and iscsi test https://review.openstack.org/508163
15:03:31 openstackgerrit Jianghua Wang proposed openstack/nova master: XenAPI: get vGPU stats from hypervisor https://review.openstack.org/512965
15:05:16 jianghuaw jaypipes, thanks for the comment. Yes, you're correct. The above is the reworked revision.
15:06:20 openstackgerrit OpenStack Proposal Bot proposed openstack/os-vif master: Updated from global requirements https://review.openstack.org/511035
15:06:22 mriedem dansmith: ok here is the bug https://bugs.launchpad.net/nova/+bug/1729371
15:06:23 openstack Launchpad bug 1729371 in OpenStack Compute (nova) "ResourceTracker races to delete instance allocations before instance is mapped to a cell" [High,Triaged]
15:06:41 dansmith mriedem: roger
15:10:35 mriedem stephenfin: can you take a look at https://review.openstack.org/#/c/481116/ and below?
15:10:52 mriedem i'd like to get that fixed and backported through all supported stable branches since it's a regression since newton
15:11:13 mriedem well, less of a regression than a busted feature since newton for evacuating with a target host
15:12:59 openstackgerrit Matthew Booth proposed openstack/nova master: Remove unused block_migration argument to _live_migration_operation https://review.openstack.org/517007
15:15:21 jianghuaw dansmith, may you take a look at this patch adding an option for enabled_vgpu_type?
15:15:23 jianghuaw https://review.openstack.org/#/c/512580/
15:17:33 dansmith jianghuaw: I have a short day today and a lot in the queue, but I will try
15:19:23 jianghuaw dansmith, understood. thanks.
15:19:53 openstackgerrit Dan Smith proposed openstack/nova master: WIP Avoid deleting allocations for instances being built https://review.openstack.org/517009
15:20:11 dansmith mriedem: this is what I'm thinking ^
15:21:32 mriedem dansmith: yeah i like that
15:21:47 mriedem still maintains the InstanceNotFound logic if the instance is in the cell db and deleted
15:21:56 dansmith we log the situation so if they see the same instance over and over, they know some allocation is stuck
15:22:20 dansmith we _could_ track that set run-to-run and log more forcefully, but that's getting more complex
15:22:43 mriedem https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcQ1b-BPsqPMmmI9aPgemul62fGgzoXzhuXWTEde_tBvLbcFaRAA
15:22:52 mriedem always good to know about the situation
15:23:01 dansmith um
15:24:29 mriedem if you want i can handle updating tests and such for that if you're busy today
15:25:22 dansmith I'm working on it now
15:25:28 dansmith I'll punt if I need to
15:30:08 mriedem anthonyper: fyi https://bugs.launchpad.net/nova/+bug/1728924
15:30:09 openstack Launchpad bug 1728924 in OpenStack Compute (nova) "console logging does not work for OL instances on xen compute" [Undecided,Confirmed]
15:32:28 dansmith mriedem: heh, we have a test that asserts that we abort these unborn children
15:32:55 mriedem i'm not sure how to appropriately respond to that statement
15:33:17 dansmith it asserts the desired behavior for reasons other than "every instance goes through this state" of course
15:33:26 dansmith but it's basically asserting that this buggy behavior exists
15:33:52 mriedem huh
15:34:00 mriedem unrelated, i suppose bauzas is in transit
15:34:12 dansmith already?
15:34:16 mriedem jaypipes: just saw another bug in triage for this same NFS resize bfv issue https://review.openstack.org/#/c/516395/
15:34:20 mriedem we should get that fixed and backported
15:34:26 mriedem s/fixed/merged/
15:34:38 mriedem note that https://review.openstack.org/#/c/516396/ verifies it
15:34:41 mriedem melwitt: you too ^
15:34:46 mriedem dansmith: only assuming
15:34:53 mriedem i know he likes to fly out about a week early :)
15:35:27 dansmith lol right
15:42:26 Nisha_Agarwal jaypipes, hi
15:42:47 Nisha_Agarwal jaypipes, i had a query regarding traits scheduling
15:44:34 Nisha_Agarwal jaypipes, is there any prefernce associated with scheduling? among traits scheduling and capabilities scheduling?
15:45:13 Nisha_Agarwal mriedem, ^^^
15:45:50 mriedem idk
15:45:58 Nisha_Agarwal johnthetubaguy, ^^^^ do u have any idea?
15:47:24 dansmith Nisha_Agarwal: what does that mean "capabilities and traits"
15:47:56 dansmith placement will consider resource amounts and traits, the scheduler will run filters on the results based on what we have today
15:49:00 Nisha_Agarwal dansmith, what do you mean by " on what we have today"?
15:49:22 Nisha_Agarwal dansmith, scheduler runs the filters after placement is executed?
15:50:11 dansmith Nisha_Agarwal: yes
15:50:39 Nisha_Agarwal dansmith, i was asking from the point that if a node has both traits and capabilities populated which one gets the preference in scheduling...is there any preference associated with it?
15:51:01 dansmith Nisha_Agarwal: and I'm saying I don't know what "capabilities" are in this context
15:51:10 mriedem dansmith: i think the CapabilitiesFilter
15:51:21 Nisha_Agarwal mriedem, yes
15:51:22 dansmith so in that case, traits win
15:51:27 dansmith as I said,
15:51:45 dansmith we select based on resource counts and traits, then apply the filters on the result
15:51:54 dansmith if we have a filter, then that will filter the result set
15:54:19 Nisha_Agarwal dansmith, but there can be a case where say the traits scheduling selects 0 nodes but capabilitiesfilter would have selected 1 node if it was run on the same original set of nodes
15:54:44 dansmith yup
15:55:47 Nisha_Agarwal dansmith, then in that case when scheduling done along with traits+capabilities, and capabilities alone differ. :(
15:56:32 dansmith Nisha_Agarwal: agree, except with the sadface
15:57:33 dansmith Nisha_Agarwal: our goal is to make scheduling better, not to preserve every detail of the existing behavior forever
15:57:41 openstackgerrit Dan Smith proposed openstack/nova master: Avoid deleting allocations for instances being built https://review.openstack.org/517009
15:57:52 dansmith mriedem: ^

Earlier   Later