| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-29 | |||
| 17:41:38 | sean-k-mooney | cdent: well its python you can always have the __init__ change then name of the class to that at runtime if your really want to punish the people debugging | |
| 17:47:58 | cdent | sean-k-mooney: punishment is the name of the game | |
| 17:48:09 | cdent | efried: yeah, that’s pretty much what I’m suggesting | |
| 17:48:58 | mriedem | we already have a thing similar to that | |
| 17:48:59 | mriedem | InstanceFaultRollback | |
| 17:49:09 | mriedem | used with a context manager | |
| 17:49:12 | mriedem | it's clear as mud | |
| 17:49:14 | efried | If you were using TaskFlow... | |
| 17:49:18 | sean-k-mooney | cdent: if the exception point new how to solve the issue though would it not do it there instead of raising | |
| 17:49:53 | cdent | sean-k-mooney: yeah, that is indeed the rub | |
| 17:49:57 | cdent | also the mudiness | |
| 17:50:51 | sean-k-mooney | cdent: it a good idea for thing that need config changes or other operator involement but it could just me in a message filed then that was logged | |
| 17:51:21 | mriedem | dansmith: you know what we need? forced host + unshelve! | |
| 17:51:23 | sean-k-mooney | for exampel. "this nova needs placement to work. go install it and read the docs" | |
| 17:51:39 | cdent | efried: the broader discussion was in this deeply branching nest of different ways in which allocatins need to be cleaned up was it better to clean up at the end of each branch or somewhere else | |
| 17:51:51 | mriedem | o/ | |
| 17:51:59 | mriedem | ~\o/~ | |
| 17:52:24 | efried | cdent Taskflow. | |
| 17:52:27 | sean-k-mooney | mriedem: so that you can unshvel an instance onto a specific host | |
| 17:52:35 | mriedem | sean-k-mooney: not only that, | |
| 17:52:38 | mriedem | but completely bypass the scheduler | |
| 17:52:44 | cdent | I was hoping that the length of the exception name was a clear indicator that I thought it was mostly crazy pants; however, all the rest of the code is already crazy pants, so who knows | |
| 17:53:14 | sean-k-mooney | mriedem: ah does the same happen with force boot to host e.g. skip scheduler | |
| 17:53:27 | mriedem | no | |
| 17:53:38 | mriedem | but it does with forced host live migration and evacuate, | |
| 17:53:42 | cdent | see above about punishment | |
| 17:53:45 | mriedem | and takashi is proposing to add the same to cold migrate | |
| 17:54:00 | mriedem | so i'm closing the loop on all the ways we can screw ourselves | |
| 17:54:57 | sean-k-mooney | well if we added it for both boot and unshivle then it would be consitet for all apis i guess. im assuming it supprot for resise also? | |
| 17:55:07 | mriedem | resize == cold migration | |
| 17:55:18 | sean-k-mooney | ah ok | |
| 17:55:23 | mriedem | sean-k-mooney: i'm sorry but i raised it up to be an asshole | |
| 17:55:28 | mriedem | you've missed that part | |
| 17:55:33 | mriedem | i don't actually want to do this | |
| 17:55:35 | mriedem | :) | |
| 17:55:39 | cdent | mriedem: help me clear up some gaps in my brains: did we generally know in advance that these edge cases were going to come up or had we thought that something would take care of it. I’m trying to grok if there’s a thing we can close up | |
| 17:56:06 | mriedem | cdent: in advance of what? placement or allowing these changes to the API to force a host and bypass the scheduler for evacuate and live migration? | |
| 17:56:15 | mriedem | s/placement/claims in the scheduler/ | |
| 17:56:15 | sean-k-mooney | haha well at least i helped make it wors by bring up the other usecases | |
| 17:56:24 | cdent | claims in the scheduler | |
| 17:56:42 | mriedem | cdent: these special move operations were not considered with claims in the scheduler at all from what i can tell | |
| 17:56:50 | mriedem | or probably move operations in general | |
| 17:56:52 | cdent | k, thanks, good datapoint | |
| 17:56:58 | mriedem | as we implemented that all as bug fixes after FF | |
| 17:57:09 | mriedem | i think, whenever we did the double up thing in the scheduler anyway | |
| 17:58:08 | mriedem | oh sorry the doubled up allocations happened the day of FF | |
| 17:58:31 | mriedem | that can't be right | |
| 17:58:52 | mriedem | oh yeah no it was FF | |
| 17:58:53 | mriedem | :( | |
| 17:59:15 | dansmith | mriedem: move ops in general yeah | |
| 18:00:02 | dansmith | I'll be pushing some more stuff up in that series in a bit | |
| 18:00:16 | dansmith | gonna try to break things into really small bits where possible per my usual, | |
| 18:00:34 | dansmith | but hopefully to make each change a clear and understandable win | |
| 18:00:42 | mriedem | ok, | |
| 18:00:51 | mriedem | i'm checking out the unshelve failure flows to see what we might have missed | |
| 18:01:00 | mriedem | and then will start working on the evacuate bug later | |
| 18:02:35 | mriedem | i guess i should start a retrospective etherpad for pike before the ptg... | |
| 18:02:44 | mriedem | i'm not sure i want to even think about what we did wrong | |
| 18:02:57 | cdent | some stuff. we admit it. done | |
| 18:02:58 | mriedem | that's reserved for when i wake up at 2am | |
| 18:11:56 | mriedem | ok https://etherpad.openstack.org/p/nova-pike-retrospective | |
| 18:11:59 | mriedem | posting to ML | |
| 18:13:10 | tomtomtom | hello, I'm having trouble launching instances with ephemeral volumes via ceph. anyone got any docs or pointers for such a configuration? | |
| 18:13:47 | tomtomtom | cinder, nova, and ceph "seem" to be creating the volume but nova comes up with "no bootable device" each time. | |
| 18:15:47 | mriedem | tomtomtom: https://docs.openstack.org/nova/latest/user/block-device-mapping.html ? | |
| 18:18:46 | sean-k-mooney | tomtomtom: does normal booting work wtih ceph backed root device or only fails with ephemeral disk | |
| 18:19:22 | mriedem | tomtomtom: can you clarify what you mean by 'ephemeral volumes'? | |
| 18:19:27 | mriedem | are you actually booting from volume? | |
| 18:19:45 | mriedem | and you consider it ephemeral because delete_on_termination=True? | |
| 18:21:06 | sean-k-mooney | alternitvly do you mean you have allocated ephmeral storage in the flavor and have confiuged nova to back all vm storage with ceph volumes | |
| 18:27:28 | mriedem | looks like we're ok wrt allocations during unshelve | |
| 18:27:39 | mriedem | conductor calls scheduler to pick a host, creates the allocations, and casts to compute | |
| 18:27:44 | mriedem | where the instance claim happens | |
| 18:27:57 | mriedem | no retries | |
| 18:28:15 | mriedem | MAYBE WE SHOULD BUILD RETRY LOGIC INTO UNSHELVE | |
| 18:29:15 | mriedem | actually, :) | |
| 18:29:26 | sean-k-mooney | cdent: lol | |
| 18:29:59 | sean-k-mooney | cdent: you realise that that would make it your problem to fix | |
| 18:30:52 | mriedem | we make the claim, and then try to spawn the instance, if that fails we unset the instance.host/node values, | |
| 18:30:57 | mriedem | but i'm not sure that we remove the allocations | |
| 18:31:50 | cdent | b@ll$ | |
| 18:32:18 | mriedem | yeah we don't | |
| 18:33:01 | mriedem | we'd abort the claim on the exit of this https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/manager.py#L4485 | |
| 18:33:22 | mriedem | https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/resource_tracker.py#L414 | |
| 18:33:33 | cdent | mriedem: maybe you need to reset your scanning algorithm: look for where we do, because we started from the point of not thinking about it, thus... | |
| 18:33:46 | mriedem | https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/resource_tracker.py#L1012 | |
| 18:33:46 | mriedem | that method, by default, sets has_ocata_computes=False | |
| 18:33:59 | mriedem | which means we won't fix the allocations https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/resource_tracker.py#L1045 | |
| 18:34:17 | mriedem | cdent: we just always relied on the periodic to heal things | |
| 18:34:29 | mriedem | which is the same code that would have done this before ^ | |
| 18:34:48 | cdent | how about we just put it back, for now? | |
| 18:35:07 | mriedem | because we want to remove it so the RT isn't trampling over thigns | |
| 18:35:23 | mriedem | it just really means that we have to be very explicity about dealing with allocations everywhere | |
| 18:35:44 | mriedem | but, that's probably for the best in the long run | |
| 18:36:12 | cdent | then in that case there’s no reason to express surprise that we aren’t handling things, yeah? | |
| 18:36:27 | mriedem | correct | |
| 18:40:00 | cdent | the functional tests are great, but fairly heavy | |
| 18:40:16 | cdent | s/heavy/cumbersome to create | |
| 18:43:37 | mriedem | https://bugs.launchpad.net/nova/+bug/1713796 | |
| 18:43:38 | openstack | Launchpad bug 1713796 in OpenStack Compute (nova) "Failed unshelve does not remove allocations from destination node" [High,Triaged] | |