Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-29
17:11:44 mriedem so it looks like we have to allocation-related bugs with evacuate,
17:11:52 mriedem 1. if you specify a host, we don't claim b/c we bypass the scheduler https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/conductor/manager.py#L749
17:12:13 mriedem 2. if you don't specify a host, we call the scheduler to find a host and create the allocations, but it rebuild fails in the compute we don't cleanup the allocations on failure
17:13:18 sean-k-mooney efried: yes but its a good start. i had a call with jay and other regarding smartnic offload of ovs where we also dissued how the dual role the whitelist is playing is problematic.
17:13:56 efried sean-k-mooney I would welcome thoughts on how the whitelist should work.
17:14:49 efried I mean, we need to have one. And it should probably be able to identify devices by classes or by specific IDs.
17:15:11 sean-k-mooney efried: basically i think it should just filter the list of pci devices and all other fuctionality such as tag or physnet association should be put into a different config option
17:16:09 efried Yeah, that makes sense.
17:16:16 sean-k-mooney device classes gets a little messy but may be usefull
17:16:31 sean-k-mooney you proably dont want to just whitelist all net devices
17:18:02 efried If it's to fit in with the other RP stuff, I would imagine the whitelist would be able to specify any of the qualitative or quantitative properties of the resource as reported by the driver.
17:18:52 efried "whitelist anything of type nic with a line speed >= 10Gbps"
17:20:29 openstackgerrit Balazs Gibizer proposed openstack/nova master: Test resource allocation during soft delete https://review.openstack.org/495159
17:20:50 openstackgerrit Elod Illes proposed openstack/nova master: Functional test: evacuate with no compute https://review.openstack.org/498482
17:21:56 efried Course, trying to figure out a way to do boolean logic in a conf var...
17:22:00 sean-k-mooney efried: the issue with all nic >=10GB is nova now needs per class special casing of the whitelist parsing
17:22:47 efried sean-k-mooney Well, the way Jay was talking about it, the whitelisting would be done by the driver.
17:22:55 efried So the driver can define whatever whitelisting syntax it wants
17:23:23 efried And it will therefore naturally be hypervisor-appropriate
17:24:11 efried That's one of the main things that I'm keeping an eye out for: that we don't wind up with a solution that works great for libvirt, but has to be kludged for e.g. HyperV or PowerVM.
17:24:13 sean-k-mooney well if we were talking about extending it to the classes as reproted but livirts nodedev-list in the libvirt dirver then that would be ok with me
17:24:23 sean-k-mooney but other wise im not sure
17:25:59 sean-k-mooney what that really means is we need to all agree on a common set of resoce clasess in placement and driver specific implentaiton in the virt dirver that convert from the hyperviros defiened view into the generic form
17:26:33 cdent queue the super upper ontology
17:28:36 sean-k-mooney well i think we can proably agree that by the time the resouce lands in the condoctor/scheduler we proably dont want to still have to care about the hyperviror on the compute node
17:29:09 efried sean-k-mooney Agree, and that's going to be a tough thing to do.
17:29:55 efried E.g. last time it was "assumed" that every PCI device would naturally have a PCI address.
17:31:45 sean-k-mooney efried: i belive that is because the pci spec requires it
17:32:12 efried whose spec?
17:32:29 mriedem cdent: new one for you https://bugs.launchpad.net/nova/+bug/1713786
17:32:30 openstack Launchpad bug 1713786 in OpenStack Compute (nova) "Allocations are not managed properly in all evacuate scenarios" [High,Triaged]
17:32:44 sean-k-mooney i belive the pci protocol is and ietf or ieee standard
17:33:01 efried Does the spec require the address to be 32 bits, and denoted as a string in domain:bus:slot.func format?
17:33:30 efried anyway, spec or no, not all hypervisors address them that way. And also, we're trying to extend to all devices, not just PCI.
17:33:55 efried which of course makes it even harder to agree on a common set of attributes for them.
17:35:35 sean-k-mooney well https://review.openstack.org/#/c/497965/2/specs/queens/approved/pci-by-device-id.rst is just for pci but we should handel other device on other busses too
17:35:46 efried Yup.
17:36:38 cdent mriedem: joy. so many twisting paths. i’m flying tomorrow and thursday so won’t personally be able to give that attention, but will try to make sure it is on the global radar
17:37:16 efried sean-k-mooney BTW, I'm not advocating for that spec to be implemented at this point. I think Jay's rebuttal is totally valid and I would love to see a more generic solution.
17:37:58 efried That said, I think if the generic solution as a baseline allowed inventorying, whitelisting, and aliasing via an opaque device ID, that would be a good starting point.
17:38:14 sean-k-mooney efried: well consideing i will need to track some nics that are not connect to the pci bus and will not have kernel netdevs in the futur more generic sound good to me
17:38:15 efried Move up to trying to create one-size-fits-all groupings/classes from there.
17:38:25 cdent mriedem: how crazy would it be to have a StuffFailedSomewhereDeleteTheAllocationsAssociatedWithConsumerContainedInThisExeption ?
17:38:49 efried cdent flake8 failed
17:39:14 cdent eagle eyes efried
17:39:30 efried Hey man, I don't make the rules. That's 86 characters.
17:40:11 cdent well crap, that kills that solution then
17:41:00 efried cdent I haven't been following the discussion, but if you're suggesting that an exception object could contain metadata that would tell the catcher how/what to clean up, I think it's a neat idea.
17:41:38 sean-k-mooney cdent: well its python you can always have the __init__ change then name of the class to that at runtime if your really want to punish the people debugging
17:47:58 cdent sean-k-mooney: punishment is the name of the game
17:48:09 cdent efried: yeah, that’s pretty much what I’m suggesting
17:48:58 mriedem we already have a thing similar to that
17:48:59 mriedem InstanceFaultRollback
17:49:09 mriedem used with a context manager
17:49:12 mriedem it's clear as mud
17:49:14 efried If you were using TaskFlow...
17:49:18 sean-k-mooney cdent: if the exception point new how to solve the issue though would it not do it there instead of raising
17:49:53 cdent sean-k-mooney: yeah, that is indeed the rub
17:49:57 cdent also the mudiness
17:50:51 sean-k-mooney cdent: it a good idea for thing that need config changes or other operator involement but it could just me in a message filed then that was logged
17:51:21 mriedem dansmith: you know what we need? forced host + unshelve!
17:51:23 sean-k-mooney for exampel. "this nova needs placement to work. go install it and read the docs"
17:51:39 cdent efried: the broader discussion was in this deeply branching nest of different ways in which allocatins need to be cleaned up was it better to clean up at the end of each branch or somewhere else
17:51:51 mriedem o/
17:51:59 mriedem ~\o/~
17:52:24 efried cdent Taskflow.
17:52:27 sean-k-mooney mriedem: so that you can unshvel an instance onto a specific host
17:52:35 mriedem sean-k-mooney: not only that,
17:52:38 mriedem but completely bypass the scheduler
17:52:44 cdent I was hoping that the length of the exception name was a clear indicator that I thought it was mostly crazy pants; however, all the rest of the code is already crazy pants, so who knows
17:53:14 sean-k-mooney mriedem: ah does the same happen with force boot to host e.g. skip scheduler
17:53:27 mriedem no
17:53:38 mriedem but it does with forced host live migration and evacuate,
17:53:42 cdent see above about punishment
17:53:45 mriedem and takashi is proposing to add the same to cold migrate
17:54:00 mriedem so i'm closing the loop on all the ways we can screw ourselves
17:54:57 sean-k-mooney well if we added it for both boot and unshivle then it would be consitet for all apis i guess. im assuming it supprot for resise also?
17:55:07 mriedem resize == cold migration
17:55:18 sean-k-mooney ah ok
17:55:23 mriedem sean-k-mooney: i'm sorry but i raised it up to be an asshole
17:55:28 mriedem you've missed that part
17:55:33 mriedem i don't actually want to do this
17:55:35 mriedem :)
17:55:39 cdent mriedem: help me clear up some gaps in my brains: did we generally know in advance that these edge cases were going to come up or had we thought that something would take care of it. I’m trying to grok if there’s a thing we can close up
17:56:06 mriedem cdent: in advance of what? placement or allowing these changes to the API to force a host and bypass the scheduler for evacuate and live migration?
17:56:15 sean-k-mooney haha well at least i helped make it wors by bring up the other usecases
17:56:15 mriedem s/placement/claims in the scheduler/
17:56:24 cdent claims in the scheduler
17:56:42 mriedem cdent: these special move operations were not considered with claims in the scheduler at all from what i can tell
17:56:50 mriedem or probably move operations in general
17:56:52 cdent k, thanks, good datapoint
17:56:58 mriedem as we implemented that all as bug fixes after FF
17:57:09 mriedem i think, whenever we did the double up thing in the scheduler anyway
17:58:08 mriedem oh sorry the doubled up allocations happened the day of FF
17:58:31 mriedem that can't be right
17:58:52 mriedem oh yeah no it was FF
17:58:53 mriedem :(
17:59:15 dansmith mriedem: move ops in general yeah
18:00:02 dansmith I'll be pushing some more stuff up in that series in a bit
18:00:16 dansmith gonna try to break things into really small bits where possible per my usual,

Earlier   Later