Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-29
17:24:11 efried That's one of the main things that I'm keeping an eye out for: that we don't wind up with a solution that works great for libvirt, but has to be kludged for e.g. HyperV or PowerVM.
17:24:13 sean-k-mooney well if we were talking about extending it to the classes as reproted but livirts nodedev-list in the libvirt dirver then that would be ok with me
17:24:23 sean-k-mooney but other wise im not sure
17:25:59 sean-k-mooney what that really means is we need to all agree on a common set of resoce clasess in placement and driver specific implentaiton in the virt dirver that convert from the hyperviros defiened view into the generic form
17:26:33 cdent queue the super upper ontology
17:28:36 sean-k-mooney well i think we can proably agree that by the time the resouce lands in the condoctor/scheduler we proably dont want to still have to care about the hyperviror on the compute node
17:29:09 efried sean-k-mooney Agree, and that's going to be a tough thing to do.
17:29:55 efried E.g. last time it was "assumed" that every PCI device would naturally have a PCI address.
17:31:45 sean-k-mooney efried: i belive that is because the pci spec requires it
17:32:12 efried whose spec?
17:32:29 mriedem cdent: new one for you https://bugs.launchpad.net/nova/+bug/1713786
17:32:30 openstack Launchpad bug 1713786 in OpenStack Compute (nova) "Allocations are not managed properly in all evacuate scenarios" [High,Triaged]
17:32:44 sean-k-mooney i belive the pci protocol is and ietf or ieee standard
17:33:01 efried Does the spec require the address to be 32 bits, and denoted as a string in domain:bus:slot.func format?
17:33:30 efried anyway, spec or no, not all hypervisors address them that way. And also, we're trying to extend to all devices, not just PCI.
17:33:55 efried which of course makes it even harder to agree on a common set of attributes for them.
17:35:35 sean-k-mooney well https://review.openstack.org/#/c/497965/2/specs/queens/approved/pci-by-device-id.rst is just for pci but we should handel other device on other busses too
17:35:46 efried Yup.
17:36:38 cdent mriedem: joy. so many twisting paths. i’m flying tomorrow and thursday so won’t personally be able to give that attention, but will try to make sure it is on the global radar
17:37:16 efried sean-k-mooney BTW, I'm not advocating for that spec to be implemented at this point. I think Jay's rebuttal is totally valid and I would love to see a more generic solution.
17:37:58 efried That said, I think if the generic solution as a baseline allowed inventorying, whitelisting, and aliasing via an opaque device ID, that would be a good starting point.
17:38:14 sean-k-mooney efried: well consideing i will need to track some nics that are not connect to the pci bus and will not have kernel netdevs in the futur more generic sound good to me
17:38:15 efried Move up to trying to create one-size-fits-all groupings/classes from there.
17:38:25 cdent mriedem: how crazy would it be to have a StuffFailedSomewhereDeleteTheAllocationsAssociatedWithConsumerContainedInThisExeption ?
17:38:49 efried cdent flake8 failed
17:39:14 cdent eagle eyes efried
17:39:30 efried Hey man, I don't make the rules. That's 86 characters.
17:40:11 cdent well crap, that kills that solution then
17:41:00 efried cdent I haven't been following the discussion, but if you're suggesting that an exception object could contain metadata that would tell the catcher how/what to clean up, I think it's a neat idea.
17:41:38 sean-k-mooney cdent: well its python you can always have the __init__ change then name of the class to that at runtime if your really want to punish the people debugging
17:47:58 cdent sean-k-mooney: punishment is the name of the game
17:48:09 cdent efried: yeah, that’s pretty much what I’m suggesting
17:48:58 mriedem we already have a thing similar to that
17:48:59 mriedem InstanceFaultRollback
17:49:09 mriedem used with a context manager
17:49:12 mriedem it's clear as mud
17:49:14 efried If you were using TaskFlow...
17:49:18 sean-k-mooney cdent: if the exception point new how to solve the issue though would it not do it there instead of raising
17:49:53 cdent sean-k-mooney: yeah, that is indeed the rub
17:49:57 cdent also the mudiness
17:50:51 sean-k-mooney cdent: it a good idea for thing that need config changes or other operator involement but it could just me in a message filed then that was logged
17:51:21 mriedem dansmith: you know what we need? forced host + unshelve!
17:51:23 sean-k-mooney for exampel. "this nova needs placement to work. go install it and read the docs"
17:51:39 cdent efried: the broader discussion was in this deeply branching nest of different ways in which allocatins need to be cleaned up was it better to clean up at the end of each branch or somewhere else
17:51:51 mriedem o/
17:51:59 mriedem ~\o/~
17:52:24 efried cdent Taskflow.
17:52:27 sean-k-mooney mriedem: so that you can unshvel an instance onto a specific host
17:52:35 mriedem sean-k-mooney: not only that,
17:52:38 mriedem but completely bypass the scheduler
17:52:44 cdent I was hoping that the length of the exception name was a clear indicator that I thought it was mostly crazy pants; however, all the rest of the code is already crazy pants, so who knows
17:53:14 sean-k-mooney mriedem: ah does the same happen with force boot to host e.g. skip scheduler
17:53:27 mriedem no
17:53:38 mriedem but it does with forced host live migration and evacuate,
17:53:42 cdent see above about punishment
17:53:45 mriedem and takashi is proposing to add the same to cold migrate
17:54:00 mriedem so i'm closing the loop on all the ways we can screw ourselves
17:54:57 sean-k-mooney well if we added it for both boot and unshivle then it would be consitet for all apis i guess. im assuming it supprot for resise also?
17:55:07 mriedem resize == cold migration
17:55:18 sean-k-mooney ah ok
17:55:23 mriedem sean-k-mooney: i'm sorry but i raised it up to be an asshole
17:55:28 mriedem you've missed that part
17:55:33 mriedem i don't actually want to do this
17:55:35 mriedem :)
17:55:39 cdent mriedem: help me clear up some gaps in my brains: did we generally know in advance that these edge cases were going to come up or had we thought that something would take care of it. I’m trying to grok if there’s a thing we can close up
17:56:06 mriedem cdent: in advance of what? placement or allowing these changes to the API to force a host and bypass the scheduler for evacuate and live migration?
17:56:15 sean-k-mooney haha well at least i helped make it wors by bring up the other usecases
17:56:15 mriedem s/placement/claims in the scheduler/
17:56:24 cdent claims in the scheduler
17:56:42 mriedem cdent: these special move operations were not considered with claims in the scheduler at all from what i can tell
17:56:50 mriedem or probably move operations in general
17:56:52 cdent k, thanks, good datapoint
17:56:58 mriedem as we implemented that all as bug fixes after FF
17:57:09 mriedem i think, whenever we did the double up thing in the scheduler anyway
17:58:08 mriedem oh sorry the doubled up allocations happened the day of FF
17:58:31 mriedem that can't be right
17:58:52 mriedem oh yeah no it was FF
17:58:53 mriedem :(
17:59:15 dansmith mriedem: move ops in general yeah
18:00:02 dansmith I'll be pushing some more stuff up in that series in a bit
18:00:16 dansmith gonna try to break things into really small bits where possible per my usual,
18:00:34 dansmith but hopefully to make each change a clear and understandable win
18:00:42 mriedem ok,
18:00:51 mriedem i'm checking out the unshelve failure flows to see what we might have missed
18:01:00 mriedem and then will start working on the evacuate bug later
18:02:35 mriedem i guess i should start a retrospective etherpad for pike before the ptg...
18:02:44 mriedem i'm not sure i want to even think about what we did wrong
18:02:57 cdent some stuff. we admit it. done
18:02:58 mriedem that's reserved for when i wake up at 2am
18:11:56 mriedem ok https://etherpad.openstack.org/p/nova-pike-retrospective
18:11:59 mriedem posting to ML
18:13:10 tomtomtom hello, I'm having trouble launching instances with ephemeral volumes via ceph. anyone got any docs or pointers for such a configuration?
18:13:47 tomtomtom cinder, nova, and ceph "seem" to be creating the volume but nova comes up with "no bootable device" each time.
18:15:47 mriedem tomtomtom: https://docs.openstack.org/nova/latest/user/block-device-mapping.html ?
18:18:46 sean-k-mooney tomtomtom: does normal booting work wtih ceph backed root device or only fails with ephemeral disk
18:19:22 mriedem tomtomtom: can you clarify what you mean by 'ephemeral volumes'?
18:19:27 mriedem are you actually booting from volume?
18:19:45 mriedem and you consider it ephemeral because delete_on_termination=True?
18:21:06 sean-k-mooney alternitvly do you mean you have allocated ephmeral storage in the flavor and have confiuged nova to back all vm storage with ceph volumes
18:27:28 mriedem looks like we're ok wrt allocations during unshelve

Earlier   Later