Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-29
17:49:18 sean-k-mooney cdent: if the exception point new how to solve the issue though would it not do it there instead of raising
17:49:53 cdent sean-k-mooney: yeah, that is indeed the rub
17:49:57 cdent also the mudiness
17:50:51 sean-k-mooney cdent: it a good idea for thing that need config changes or other operator involement but it could just me in a message filed then that was logged
17:51:21 mriedem dansmith: you know what we need? forced host + unshelve!
17:51:23 sean-k-mooney for exampel. "this nova needs placement to work. go install it and read the docs"
17:51:39 cdent efried: the broader discussion was in this deeply branching nest of different ways in which allocatins need to be cleaned up was it better to clean up at the end of each branch or somewhere else
17:51:51 mriedem o/
17:51:59 mriedem ~\o/~
17:52:24 efried cdent Taskflow.
17:52:27 sean-k-mooney mriedem: so that you can unshvel an instance onto a specific host
17:52:35 mriedem sean-k-mooney: not only that,
17:52:38 mriedem but completely bypass the scheduler
17:52:44 cdent I was hoping that the length of the exception name was a clear indicator that I thought it was mostly crazy pants; however, all the rest of the code is already crazy pants, so who knows
17:53:14 sean-k-mooney mriedem: ah does the same happen with force boot to host e.g. skip scheduler
17:53:27 mriedem no
17:53:38 mriedem but it does with forced host live migration and evacuate,
17:53:42 cdent see above about punishment
17:53:45 mriedem and takashi is proposing to add the same to cold migrate
17:54:00 mriedem so i'm closing the loop on all the ways we can screw ourselves
17:54:57 sean-k-mooney well if we added it for both boot and unshivle then it would be consitet for all apis i guess. im assuming it supprot for resise also?
17:55:07 mriedem resize == cold migration
17:55:18 sean-k-mooney ah ok
17:55:23 mriedem sean-k-mooney: i'm sorry but i raised it up to be an asshole
17:55:28 mriedem you've missed that part
17:55:33 mriedem i don't actually want to do this
17:55:35 mriedem :)
17:55:39 cdent mriedem: help me clear up some gaps in my brains: did we generally know in advance that these edge cases were going to come up or had we thought that something would take care of it. I’m trying to grok if there’s a thing we can close up
17:56:06 mriedem cdent: in advance of what? placement or allowing these changes to the API to force a host and bypass the scheduler for evacuate and live migration?
17:56:15 sean-k-mooney haha well at least i helped make it wors by bring up the other usecases
17:56:15 mriedem s/placement/claims in the scheduler/
17:56:24 cdent claims in the scheduler
17:56:42 mriedem cdent: these special move operations were not considered with claims in the scheduler at all from what i can tell
17:56:50 mriedem or probably move operations in general
17:56:52 cdent k, thanks, good datapoint
17:56:58 mriedem as we implemented that all as bug fixes after FF
17:57:09 mriedem i think, whenever we did the double up thing in the scheduler anyway
17:58:08 mriedem oh sorry the doubled up allocations happened the day of FF
17:58:31 mriedem that can't be right
17:58:52 mriedem oh yeah no it was FF
17:58:53 mriedem :(
17:59:15 dansmith mriedem: move ops in general yeah
18:00:02 dansmith I'll be pushing some more stuff up in that series in a bit
18:00:16 dansmith gonna try to break things into really small bits where possible per my usual,
18:00:34 dansmith but hopefully to make each change a clear and understandable win
18:00:42 mriedem ok,
18:00:51 mriedem i'm checking out the unshelve failure flows to see what we might have missed
18:01:00 mriedem and then will start working on the evacuate bug later
18:02:35 mriedem i guess i should start a retrospective etherpad for pike before the ptg...
18:02:44 mriedem i'm not sure i want to even think about what we did wrong
18:02:57 cdent some stuff. we admit it. done
18:02:58 mriedem that's reserved for when i wake up at 2am
18:11:56 mriedem ok https://etherpad.openstack.org/p/nova-pike-retrospective
18:11:59 mriedem posting to ML
18:13:10 tomtomtom hello, I'm having trouble launching instances with ephemeral volumes via ceph. anyone got any docs or pointers for such a configuration?
18:13:47 tomtomtom cinder, nova, and ceph "seem" to be creating the volume but nova comes up with "no bootable device" each time.
18:15:47 mriedem tomtomtom: https://docs.openstack.org/nova/latest/user/block-device-mapping.html ?
18:18:46 sean-k-mooney tomtomtom: does normal booting work wtih ceph backed root device or only fails with ephemeral disk
18:19:22 mriedem tomtomtom: can you clarify what you mean by 'ephemeral volumes'?
18:19:27 mriedem are you actually booting from volume?
18:19:45 mriedem and you consider it ephemeral because delete_on_termination=True?
18:21:06 sean-k-mooney alternitvly do you mean you have allocated ephmeral storage in the flavor and have confiuged nova to back all vm storage with ceph volumes
18:27:28 mriedem looks like we're ok wrt allocations during unshelve
18:27:39 mriedem conductor calls scheduler to pick a host, creates the allocations, and casts to compute
18:27:44 mriedem where the instance claim happens
18:27:57 mriedem no retries
18:28:15 mriedem MAYBE WE SHOULD BUILD RETRY LOGIC INTO UNSHELVE
18:29:15 mriedem actually, :)
18:29:26 sean-k-mooney cdent: lol
18:29:59 sean-k-mooney cdent: you realise that that would make it your problem to fix
18:30:52 mriedem we make the claim, and then try to spawn the instance, if that fails we unset the instance.host/node values,
18:30:57 mriedem but i'm not sure that we remove the allocations
18:31:50 cdent b@ll$
18:32:18 mriedem yeah we don't
18:33:01 mriedem we'd abort the claim on the exit of this https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/manager.py#L4485
18:33:22 mriedem https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/resource_tracker.py#L414
18:33:33 cdent mriedem: maybe you need to reset your scanning algorithm: look for where we do, because we started from the point of not thinking about it, thus...
18:33:46 mriedem that method, by default, sets has_ocata_computes=False
18:33:46 mriedem https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/resource_tracker.py#L1012
18:33:59 mriedem which means we won't fix the allocations https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/resource_tracker.py#L1045
18:34:17 mriedem cdent: we just always relied on the periodic to heal things
18:34:29 mriedem which is the same code that would have done this before ^
18:34:48 cdent how about we just put it back, for now?
18:35:07 mriedem because we want to remove it so the RT isn't trampling over thigns
18:35:23 mriedem it just really means that we have to be very explicity about dealing with allocations everywhere
18:35:44 mriedem but, that's probably for the best in the long run
18:36:12 cdent then in that case there’s no reason to express surprise that we aren’t handling things, yeah?
18:36:27 mriedem correct
18:40:00 cdent the functional tests are great, but fairly heavy
18:40:16 cdent s/heavy/cumbersome to create
18:43:37 mriedem https://bugs.launchpad.net/nova/+bug/1713796
18:43:38 openstack Launchpad bug 1713796 in OpenStack Compute (nova) "Failed unshelve does not remove allocations from destination node" [High,Triaged]
18:43:51 mriedem there are at least 2 ways unshelve can fail there which we don't cleanup the allocations
18:45:03 mriedem maybe we need an undo_allocations decorator for several methods in the compute manager
18:46:41 cdent are the exit conditions workable for a decorator (or, to put it another way, how is that different from the two other ideas above?)
18:47:32 mriedem it could be messy in a decorator, probably lots of conditional logic based on the operation being performed
18:47:37 mriedem which is based on the task_state
18:48:03 mriedem so i'm not going to bother thinking about that for now
18:49:06 mriedem regarding the functional tests, i think we need those regardless
18:49:16 cdent oh, yeah, I wasn’t saying we should get rid of them

Earlier   Later