| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-29 | |||
| 16:19:17 | mriedem | since the docs migration to use the new sphinx theme | |
| 16:19:18 | mriedem | RAR | |
| 16:19:37 | mriedem | ssmith: openstack help compute or something should give you the optoins | |
| 16:19:48 | mriedem | it's --compute-api-version or something | |
| 16:22:22 | mriedem | ssmith: --os-compute-api-version i think | |
| 16:22:32 | ssmith | ok | |
| 16:23:08 | mriedem | i don't know if the osc cli whitelists the fields it shows though | |
| 16:23:41 | ssmith | Compute API version, default=2.1 | |
| 16:23:41 | ssmith | os-compute-api-version <compute-api-version> | |
| 16:24:07 | mriedem | yeah, so specify 2.9 | |
| 16:25:04 | ssmith | openstack server show a29d1e77-5191-4629-a913-ae98ed22e284 --os-compute-api-version 2.9 WORKED | |
| 16:25:16 | mriedem | cool | |
| 16:25:41 | ssmith | Any way to permanently set the version? | |
| 16:25:52 | mriedem | env var | |
| 16:25:59 | mriedem | OS_COMPUTE_API_VERSION=2.9 i think | |
| 16:28:39 | ssmith | Did a "export OS_COMPUTE_API_VERSION=2.9" which worked | |
| 16:31:07 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Cleanup allocations on invalid dest node during live migration https://review.openstack.org/498861 | |
| 16:31:34 | mriedem | dansmith: cdent: gibi: here is the fix for the live migration pre-check error bug ^ | |
| 16:33:16 | efried | sean-k-mooney that first bp was https://blueprints.launchpad.net/nova/+spec/pci-by-device-id whose spec (https://review.openstack.org/497965) has some good discussion started. | |
| 16:34:16 | efried | jaypipes and I talked about it yesterday a bit and he encouraged me to move some of the main themes over to the other one stephenfin mentioned - https://blueprints.launchpad.net/nova/+spec/devices-as-resources - whose spec seed (not yet started) is here: https://review.openstack.org/497978 | |
| 16:34:27 | efried | ...and to start an etherpad for same topic for the PTG. | |
| 16:37:14 | openstackgerrit | Lucian Petrut proposed openstack/nova master: Fix nova assisted volume snapshots https://review.openstack.org/498845 | |
| 16:37:36 | openstackgerrit | Lajos Katona proposed openstack/nova master: WIP: Test server movings with custom resources https://review.openstack.org/497399 | |
| 16:41:27 | dansmith | mriedem: so, my feeling as we neared the end of pike was that we were really not doing ourselves any favors by trying to put the allocation stuff into the existing RT calls since we need to do different things and need some context from the compute manager to do the right thing | |
| 16:41:43 | dansmith | which is why we do some silly stuff in RT like looking at if prefix == 'old_' to do certain things | |
| 16:42:13 | dansmith | mriedem: for this migration uuid thing, I kinda want to add these new paths to compute manager itself, so at least we can delete RT eventually without needing to move things out of it | |
| 16:42:18 | dansmith | does that sound legit? | |
| 16:42:38 | dansmith | it might mean moving some of our existing allocation handling back out of RT as well and just clearly marking which bits are legacy pike behavior and not | |
| 16:46:00 | cdent | dansmith: I can only speak for myself, but I think that’s totally legit | |
| 16:47:13 | mriedem | dansmith: i think in general that's OK given we've already started duplicating some allocation-specific stuff in the compute manager outside of the RT | |
| 16:47:28 | dansmith | yeah | |
| 16:47:39 | mriedem | e.g. https://review.openstack.org/#/c/496976/ | |
| 16:47:53 | dansmith | I just spent 45 minutes trying to detect the condition I need from down in RT and it just doesn't make any sense I think | |
| 16:48:05 | mriedem | it's also harder to debug | |
| 16:48:08 | mriedem | because of the layering | |
| 16:48:15 | dansmith | yeah | |
| 16:48:22 | mriedem | RT is always a new journey for me everytime i have to look into it | |
| 16:59:27 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Refactor LiveMigrationTask._find_destination https://review.openstack.org/498874 | |
| 17:05:43 | sean-k-mooney | efried: this one https://review.openstack.org/#/c/497965/2/specs/queens/approved/pci-by-device-id.rst | |
| 17:05:54 | sean-k-mooney | ill take a look at it | |
| 17:10:28 | efried | sean-k-mooney Great, thanks. FYI that one's not going to fly as currently conceived, but it's got some good problem descriptions and the discussion is getting us moving in the right direction. | |
| 17:10:44 | openstackgerrit | Merged openstack/nova master: Enhance support matrix document https://review.openstack.org/482013 | |
| 17:11:44 | mriedem | so it looks like we have to allocation-related bugs with evacuate, | |
| 17:11:52 | mriedem | 1. if you specify a host, we don't claim b/c we bypass the scheduler https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/conductor/manager.py#L749 | |
| 17:12:13 | mriedem | 2. if you don't specify a host, we call the scheduler to find a host and create the allocations, but it rebuild fails in the compute we don't cleanup the allocations on failure | |
| 17:13:18 | sean-k-mooney | efried: yes but its a good start. i had a call with jay and other regarding smartnic offload of ovs where we also dissued how the dual role the whitelist is playing is problematic. | |
| 17:13:56 | efried | sean-k-mooney I would welcome thoughts on how the whitelist should work. | |
| 17:14:49 | efried | I mean, we need to have one. And it should probably be able to identify devices by classes or by specific IDs. | |
| 17:15:11 | sean-k-mooney | efried: basically i think it should just filter the list of pci devices and all other fuctionality such as tag or physnet association should be put into a different config option | |
| 17:16:09 | efried | Yeah, that makes sense. | |
| 17:16:16 | sean-k-mooney | device classes gets a little messy but may be usefull | |
| 17:16:31 | sean-k-mooney | you proably dont want to just whitelist all net devices | |
| 17:18:02 | efried | If it's to fit in with the other RP stuff, I would imagine the whitelist would be able to specify any of the qualitative or quantitative properties of the resource as reported by the driver. | |
| 17:18:52 | efried | "whitelist anything of type nic with a line speed >= 10Gbps" | |
| 17:20:29 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Test resource allocation during soft delete https://review.openstack.org/495159 | |
| 17:20:50 | openstackgerrit | Elod Illes proposed openstack/nova master: Functional test: evacuate with no compute https://review.openstack.org/498482 | |
| 17:21:56 | efried | Course, trying to figure out a way to do boolean logic in a conf var... | |
| 17:22:00 | sean-k-mooney | efried: the issue with all nic >=10GB is nova now needs per class special casing of the whitelist parsing | |
| 17:22:47 | efried | sean-k-mooney Well, the way Jay was talking about it, the whitelisting would be done by the driver. | |
| 17:22:55 | efried | So the driver can define whatever whitelisting syntax it wants | |
| 17:23:23 | efried | And it will therefore naturally be hypervisor-appropriate | |
| 17:24:11 | efried | That's one of the main things that I'm keeping an eye out for: that we don't wind up with a solution that works great for libvirt, but has to be kludged for e.g. HyperV or PowerVM. | |
| 17:24:13 | sean-k-mooney | well if we were talking about extending it to the classes as reproted but livirts nodedev-list in the libvirt dirver then that would be ok with me | |
| 17:24:23 | sean-k-mooney | but other wise im not sure | |
| 17:25:59 | sean-k-mooney | what that really means is we need to all agree on a common set of resoce clasess in placement and driver specific implentaiton in the virt dirver that convert from the hyperviros defiened view into the generic form | |
| 17:26:33 | cdent | queue the super upper ontology | |
| 17:28:36 | sean-k-mooney | well i think we can proably agree that by the time the resouce lands in the condoctor/scheduler we proably dont want to still have to care about the hyperviror on the compute node | |
| 17:29:09 | efried | sean-k-mooney Agree, and that's going to be a tough thing to do. | |
| 17:29:55 | efried | E.g. last time it was "assumed" that every PCI device would naturally have a PCI address. | |
| 17:31:45 | sean-k-mooney | efried: i belive that is because the pci spec requires it | |
| 17:32:12 | efried | whose spec? | |
| 17:32:29 | mriedem | cdent: new one for you https://bugs.launchpad.net/nova/+bug/1713786 | |
| 17:32:30 | openstack | Launchpad bug 1713786 in OpenStack Compute (nova) "Allocations are not managed properly in all evacuate scenarios" [High,Triaged] | |
| 17:32:44 | sean-k-mooney | i belive the pci protocol is and ietf or ieee standard | |
| 17:33:01 | efried | Does the spec require the address to be 32 bits, and denoted as a string in domain:bus:slot.func format? | |
| 17:33:30 | efried | anyway, spec or no, not all hypervisors address them that way. And also, we're trying to extend to all devices, not just PCI. | |
| 17:33:55 | efried | which of course makes it even harder to agree on a common set of attributes for them. | |
| 17:35:35 | sean-k-mooney | well https://review.openstack.org/#/c/497965/2/specs/queens/approved/pci-by-device-id.rst is just for pci but we should handel other device on other busses too | |
| 17:35:46 | efried | Yup. | |
| 17:36:38 | cdent | mriedem: joy. so many twisting paths. i’m flying tomorrow and thursday so won’t personally be able to give that attention, but will try to make sure it is on the global radar | |
| 17:37:16 | efried | sean-k-mooney BTW, I'm not advocating for that spec to be implemented at this point. I think Jay's rebuttal is totally valid and I would love to see a more generic solution. | |
| 17:37:58 | efried | That said, I think if the generic solution as a baseline allowed inventorying, whitelisting, and aliasing via an opaque device ID, that would be a good starting point. | |
| 17:38:14 | sean-k-mooney | efried: well consideing i will need to track some nics that are not connect to the pci bus and will not have kernel netdevs in the futur more generic sound good to me | |
| 17:38:15 | efried | Move up to trying to create one-size-fits-all groupings/classes from there. | |
| 17:38:25 | cdent | mriedem: how crazy would it be to have a StuffFailedSomewhereDeleteTheAllocationsAssociatedWithConsumerContainedInThisExeption ? | |
| 17:38:49 | efried | cdent flake8 failed | |
| 17:39:14 | cdent | eagle eyes efried | |
| 17:39:30 | efried | Hey man, I don't make the rules. That's 86 characters. | |
| 17:40:11 | cdent | well crap, that kills that solution then | |
| 17:41:00 | efried | cdent I haven't been following the discussion, but if you're suggesting that an exception object could contain metadata that would tell the catcher how/what to clean up, I think it's a neat idea. | |
| 17:41:38 | sean-k-mooney | cdent: well its python you can always have the __init__ change then name of the class to that at runtime if your really want to punish the people debugging | |
| 17:47:58 | cdent | sean-k-mooney: punishment is the name of the game | |
| 17:48:09 | cdent | efried: yeah, that’s pretty much what I’m suggesting | |
| 17:48:58 | mriedem | we already have a thing similar to that | |
| 17:48:59 | mriedem | InstanceFaultRollback | |
| 17:49:09 | mriedem | used with a context manager | |
| 17:49:12 | mriedem | it's clear as mud | |
| 17:49:14 | efried | If you were using TaskFlow... | |
| 17:49:18 | sean-k-mooney | cdent: if the exception point new how to solve the issue though would it not do it there instead of raising | |
| 17:49:53 | cdent | sean-k-mooney: yeah, that is indeed the rub | |