| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-29 | |||
| 17:32:44 | sean-k-mooney | i belive the pci protocol is and ietf or ieee standard | |
| 17:33:01 | efried | Does the spec require the address to be 32 bits, and denoted as a string in domain:bus:slot.func format? | |
| 17:33:30 | efried | anyway, spec or no, not all hypervisors address them that way. And also, we're trying to extend to all devices, not just PCI. | |
| 17:33:55 | efried | which of course makes it even harder to agree on a common set of attributes for them. | |
| 17:35:35 | sean-k-mooney | well https://review.openstack.org/#/c/497965/2/specs/queens/approved/pci-by-device-id.rst is just for pci but we should handel other device on other busses too | |
| 17:35:46 | efried | Yup. | |
| 17:36:38 | cdent | mriedem: joy. so many twisting paths. i’m flying tomorrow and thursday so won’t personally be able to give that attention, but will try to make sure it is on the global radar | |
| 17:37:16 | efried | sean-k-mooney BTW, I'm not advocating for that spec to be implemented at this point. I think Jay's rebuttal is totally valid and I would love to see a more generic solution. | |
| 17:37:58 | efried | That said, I think if the generic solution as a baseline allowed inventorying, whitelisting, and aliasing via an opaque device ID, that would be a good starting point. | |
| 17:38:14 | sean-k-mooney | efried: well consideing i will need to track some nics that are not connect to the pci bus and will not have kernel netdevs in the futur more generic sound good to me | |
| 17:38:15 | efried | Move up to trying to create one-size-fits-all groupings/classes from there. | |
| 17:38:25 | cdent | mriedem: how crazy would it be to have a StuffFailedSomewhereDeleteTheAllocationsAssociatedWithConsumerContainedInThisExeption ? | |
| 17:38:49 | efried | cdent flake8 failed | |
| 17:39:14 | cdent | eagle eyes efried | |
| 17:39:30 | efried | Hey man, I don't make the rules. That's 86 characters. | |
| 17:40:11 | cdent | well crap, that kills that solution then | |
| 17:41:00 | efried | cdent I haven't been following the discussion, but if you're suggesting that an exception object could contain metadata that would tell the catcher how/what to clean up, I think it's a neat idea. | |
| 17:41:38 | sean-k-mooney | cdent: well its python you can always have the __init__ change then name of the class to that at runtime if your really want to punish the people debugging | |
| 17:47:58 | cdent | sean-k-mooney: punishment is the name of the game | |
| 17:48:09 | cdent | efried: yeah, that’s pretty much what I’m suggesting | |
| 17:48:58 | mriedem | we already have a thing similar to that | |
| 17:48:59 | mriedem | InstanceFaultRollback | |
| 17:49:09 | mriedem | used with a context manager | |
| 17:49:12 | mriedem | it's clear as mud | |
| 17:49:14 | efried | If you were using TaskFlow... | |
| 17:49:18 | sean-k-mooney | cdent: if the exception point new how to solve the issue though would it not do it there instead of raising | |
| 17:49:53 | cdent | sean-k-mooney: yeah, that is indeed the rub | |
| 17:49:57 | cdent | also the mudiness | |
| 17:50:51 | sean-k-mooney | cdent: it a good idea for thing that need config changes or other operator involement but it could just me in a message filed then that was logged | |
| 17:51:21 | mriedem | dansmith: you know what we need? forced host + unshelve! | |
| 17:51:23 | sean-k-mooney | for exampel. "this nova needs placement to work. go install it and read the docs" | |
| 17:51:39 | cdent | efried: the broader discussion was in this deeply branching nest of different ways in which allocatins need to be cleaned up was it better to clean up at the end of each branch or somewhere else | |
| 17:51:51 | mriedem | o/ | |
| 17:51:59 | mriedem | ~\o/~ | |
| 17:52:24 | efried | cdent Taskflow. | |
| 17:52:27 | sean-k-mooney | mriedem: so that you can unshvel an instance onto a specific host | |
| 17:52:35 | mriedem | sean-k-mooney: not only that, | |
| 17:52:38 | mriedem | but completely bypass the scheduler | |
| 17:52:44 | cdent | I was hoping that the length of the exception name was a clear indicator that I thought it was mostly crazy pants; however, all the rest of the code is already crazy pants, so who knows | |
| 17:53:14 | sean-k-mooney | mriedem: ah does the same happen with force boot to host e.g. skip scheduler | |
| 17:53:27 | mriedem | no | |
| 17:53:38 | mriedem | but it does with forced host live migration and evacuate, | |
| 17:53:42 | cdent | see above about punishment | |
| 17:53:45 | mriedem | and takashi is proposing to add the same to cold migrate | |
| 17:54:00 | mriedem | so i'm closing the loop on all the ways we can screw ourselves | |
| 17:54:57 | sean-k-mooney | well if we added it for both boot and unshivle then it would be consitet for all apis i guess. im assuming it supprot for resise also? | |
| 17:55:07 | mriedem | resize == cold migration | |
| 17:55:18 | sean-k-mooney | ah ok | |
| 17:55:23 | mriedem | sean-k-mooney: i'm sorry but i raised it up to be an asshole | |
| 17:55:28 | mriedem | you've missed that part | |
| 17:55:33 | mriedem | i don't actually want to do this | |
| 17:55:35 | mriedem | :) | |
| 17:55:39 | cdent | mriedem: help me clear up some gaps in my brains: did we generally know in advance that these edge cases were going to come up or had we thought that something would take care of it. I’m trying to grok if there’s a thing we can close up | |
| 17:56:06 | mriedem | cdent: in advance of what? placement or allowing these changes to the API to force a host and bypass the scheduler for evacuate and live migration? | |
| 17:56:15 | sean-k-mooney | haha well at least i helped make it wors by bring up the other usecases | |
| 17:56:15 | mriedem | s/placement/claims in the scheduler/ | |
| 17:56:24 | cdent | claims in the scheduler | |
| 17:56:42 | mriedem | cdent: these special move operations were not considered with claims in the scheduler at all from what i can tell | |
| 17:56:50 | mriedem | or probably move operations in general | |
| 17:56:52 | cdent | k, thanks, good datapoint | |
| 17:56:58 | mriedem | as we implemented that all as bug fixes after FF | |
| 17:57:09 | mriedem | i think, whenever we did the double up thing in the scheduler anyway | |
| 17:58:08 | mriedem | oh sorry the doubled up allocations happened the day of FF | |
| 17:58:31 | mriedem | that can't be right | |
| 17:58:52 | mriedem | oh yeah no it was FF | |
| 17:58:53 | mriedem | :( | |
| 17:59:15 | dansmith | mriedem: move ops in general yeah | |
| 18:00:02 | dansmith | I'll be pushing some more stuff up in that series in a bit | |
| 18:00:16 | dansmith | gonna try to break things into really small bits where possible per my usual, | |
| 18:00:34 | dansmith | but hopefully to make each change a clear and understandable win | |
| 18:00:42 | mriedem | ok, | |
| 18:00:51 | mriedem | i'm checking out the unshelve failure flows to see what we might have missed | |
| 18:01:00 | mriedem | and then will start working on the evacuate bug later | |
| 18:02:35 | mriedem | i guess i should start a retrospective etherpad for pike before the ptg... | |
| 18:02:44 | mriedem | i'm not sure i want to even think about what we did wrong | |
| 18:02:57 | cdent | some stuff. we admit it. done | |
| 18:02:58 | mriedem | that's reserved for when i wake up at 2am | |
| 18:11:56 | mriedem | ok https://etherpad.openstack.org/p/nova-pike-retrospective | |
| 18:11:59 | mriedem | posting to ML | |
| 18:13:10 | tomtomtom | hello, I'm having trouble launching instances with ephemeral volumes via ceph. anyone got any docs or pointers for such a configuration? | |
| 18:13:47 | tomtomtom | cinder, nova, and ceph "seem" to be creating the volume but nova comes up with "no bootable device" each time. | |
| 18:15:47 | mriedem | tomtomtom: https://docs.openstack.org/nova/latest/user/block-device-mapping.html ? | |
| 18:18:46 | sean-k-mooney | tomtomtom: does normal booting work wtih ceph backed root device or only fails with ephemeral disk | |
| 18:19:22 | mriedem | tomtomtom: can you clarify what you mean by 'ephemeral volumes'? | |
| 18:19:27 | mriedem | are you actually booting from volume? | |
| 18:19:45 | mriedem | and you consider it ephemeral because delete_on_termination=True? | |
| 18:21:06 | sean-k-mooney | alternitvly do you mean you have allocated ephmeral storage in the flavor and have confiuged nova to back all vm storage with ceph volumes | |
| 18:27:28 | mriedem | looks like we're ok wrt allocations during unshelve | |
| 18:27:39 | mriedem | conductor calls scheduler to pick a host, creates the allocations, and casts to compute | |
| 18:27:44 | mriedem | where the instance claim happens | |
| 18:27:57 | mriedem | no retries | |
| 18:28:15 | mriedem | MAYBE WE SHOULD BUILD RETRY LOGIC INTO UNSHELVE | |
| 18:29:15 | mriedem | actually, :) | |
| 18:29:26 | sean-k-mooney | cdent: lol | |
| 18:29:59 | sean-k-mooney | cdent: you realise that that would make it your problem to fix | |
| 18:30:52 | mriedem | we make the claim, and then try to spawn the instance, if that fails we unset the instance.host/node values, | |
| 18:30:57 | mriedem | but i'm not sure that we remove the allocations | |
| 18:31:50 | cdent | b@ll$ | |
| 18:32:18 | mriedem | yeah we don't | |
| 18:33:01 | mriedem | we'd abort the claim on the exit of this https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/manager.py#L4485 | |