Earlier  
Posted Nick Remark
#openstack-nova - 2017-07-31
14:54:38 gibi_ is there a limit on paste.openstack.org about the size of the log?
14:55:35 clarkb gibi_: yes its like 1MB or something. If you have to paste very large amounts of text I think gist.github.com allows large pastes anonymously
14:55:46 gibi_ clarkb: thanks
14:56:00 gibi_ it seems pastebin doesn't have such limit either
14:56:01 gibi_ https://pastebin.com/0DuaUrZJ
14:56:25 sdague mdbooth: looking
14:56:31 sdague bauzas: thanks!
14:57:08 bauzas I usually decrease by 5/6 bugs per day
14:57:22 bauzas given we had 120 new ones...
14:57:31 bauzas so, yeah, very impressive
14:58:17 bauzas mriedem: had a chance to qualify the pike-rc-candidates ?
14:58:20 bauzas I gave you 3 of those
14:58:33 bauzas mriedem: I can look over the rest
14:58:52 dansmith jaypipes: cool, that looks like it'll tell us what is happeing
14:59:23 mriedem bauzas: no
14:59:26 jaypipes dansmith: cool. FYI, bhagyashris is also functionally testing shared storage with NFS and the claims-in-scheduler patch.
14:59:35 bauzas mriedem: okay, will review those
14:59:50 bauzas also, I need to make sure we have the claims bugs in there ^
15:00:23 mriedem i marked both claims bugs for rc
15:00:41 openstackgerrit Spencer Yu proposed openstack/python-novaclient master: Nova client should retry with Retry-After value https://review.openstack.org/447766
15:01:21 openstack Launchpad bug 1707256 in OpenStack Compute (nova) "Scheduler report client does not account for shared resource providers" [High,Confirmed] - Assigned to Jay Pipes (jaypipes)
15:01:21 mriedem https://bugs.launchpad.net/nova/+bug/1707256
15:01:25 openstack Launchpad bug 1707252 in OpenStack Compute (nova) "Claims in the scheduler does not account for doubling allocations on resize to same host" [Medium,Confirmed]
15:01:25 mriedem https://bugs.launchpad.net/nova/+bug/1707252
15:01:30 bauzas ack
15:01:31 mriedem are the 2 i created on friday
15:01:46 gibi_ jaypipes, dansmith, cdent: relevant part of the log is here http://paste.openstack.org/show/617028/
15:04:11 gibi_ it seems that the code correctly sends the stripped allocations
15:04:46 gibi_ but after that PUT placement still has the old allocation as well
15:05:47 gibi_ does PUT /placement/allocations expected to totally overwrite the db for the instance
15:05:50 gibi_ ?
15:06:04 melwitt dansmith, mriedem: I have a fix up for a volume detach data corruption bug at https://review.openstack.org/#/c/488545/ that was caused by an earlier attempt to fix a different bug. has to be backported all the way to newton I think
15:06:05 mriedem yes
15:06:07 mriedem gibi_: yes
15:06:49 mriedem melwitt: good lord
15:07:05 mriedem i don't think the backports to newton ever landed because i also depended on them for another series
15:07:23 melwitt o rly
15:07:33 mriedem oh nvm https://review.openstack.org/#/c/425114/
15:07:38 mriedem must be something else then
15:07:56 cdent gibi_: is there yet another PUT after the stripped one?
15:07:59 mriedem i was thinking of this series i have in newton https://review.openstack.org/#/c/470347/
15:08:05 mriedem to wait for an interface to be detached
15:08:15 melwitt oh, okay
15:08:39 cdent gibi_: or is maybe the one with the stripped not being accepted (because of 409)?
15:09:57 gibi_ cdent: look at line 5-7 in http://paste.openstack.org/show/617028/
15:10:07 gibi_ cdent: sorry 4-7
15:10:07 cdent yeah, I'm there now
15:10:19 gibi_ cdent: 4 sends an allocation list with one item
15:10:52 gibi_ but line 6 writes two allocations to the db
15:12:01 gibi_ cdent: not two, three actually
15:12:12 mriedem melwitt: i think i was thinking of this https://review.openstack.org/#/c/441204/
15:12:16 mriedem which is part of that newton series
15:12:20 mriedem and sounds similar to what you're doing
15:12:22 gibi_ cdent: but all for the same provider so this is not the problem
15:13:10 cdent gibi_: right, what you are seeing there is just an artifact of the object: an REST-level allocation is made up of multiple Allocation Objects
15:13:31 gibi_ cdent: yeah, I see now. sorry
15:13:41 melwitt mriedem: oh, right. I noticed that too at some point thinking it's related but it seems like it's not, i.e. in the case of the bug I think isPersistent() will still return True
15:13:42 gibi_ cdent: anyhow there is no other PUT on allocations later in the log
15:13:53 cdent yeah,
15:14:52 melwitt mriedem: the domain will still be persistent even if the volume had been detached from the persistent config in the past
15:15:01 cdent gibi_: line 148 is demonstrating how things are wrong, correct? that's after confirm, and some time to let things settled?
15:15:38 cdent so source hasn't cleaned up
15:16:09 gibi_ cdent: yes
15:17:49 cdent gibi_: so the problem is somewhere near here: https://review.openstack.org/#/c/488510/4/nova/compute/resource_tracker.py@1083
15:18:16 cdent sorry, not there
15:18:36 cdent 476
15:19:33 gibi_ but that is the piece of code that generates our PUT
15:19:40 gibi_ to me
15:19:40 gibi_ and that PUT seems correct
15:19:45 cdent true
15:26:47 gibi_ I added code to https://review.openstack.org/#/c/488510/4/nova/scheduler/client/report.py@1079 to read back the allocations the code just PUT-ed
15:26:54 gibi_ http://paste.openstack.org/show/617031/
15:27:15 gibi_ and the GET returns both allocations after the PUT
15:27:31 gibi_ so the problem is in the placement I think
15:28:04 cdent wow
15:28:15 cdent that will be an exciting bug if so
15:28:46 dansmith I'm missing the obvious thing
15:29:00 cdent gibi_: yeah
15:29:03 dansmith oh, we can't read back the allocations we just wrote in a particular place?
15:29:17 cdent it only deletes where rp uuid and consume uuid ==
15:29:21 cdent not just consumer uuid
15:29:47 bauzas interesting
15:30:16 cdent https://github.com/openstack/nova/blob/master/nova/objects/resource_provider.py#L1509-L1520
15:30:43 cdent gibi_: see what happens if you get rid of the first == in the and_
15:30:52 gibi_ checking
15:31:01 bauzas yeah that will work
15:31:09 bauzas obviously
15:31:17 cdent dansmith: we had in our brains that we were removing all allocations for the consumer
15:31:24 cdent but the code is only removing some of them
15:32:33 dansmith meaning, the api is such that all should be removed, but some bug in the db side of placement prevents that from happening?
15:32:50 cdent yeah
15:32:59 dansmith that seems ungood
15:33:15 cdent it looks like the original goal was to _only_ remove exact matches (not sure why)
15:33:22 cdent (the comment says as much)
15:33:27 dansmith hmm
15:33:41 dansmith jaypipes seemed to think that should be fully atomic, so that seems weird
15:33:59 gibi_ I can confirm that removing the rp == from the db code makes the placement API behave as expected in this particular case http://paste.openstack.org/show/617034/
15:34:16 cdent and we had three authors on that particular change, so is probably going to be hard to remember the whys and wherefores
15:34:18 cdent gibi_: nice
15:34:56 cdent dansmith: since we've been assuming all this time that the behavior is one thing and not the other, I think we should just change it

Earlier   Later