Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-27
13:14:05 cdent efried: So, short answer is “yes” so second question is “Do you yet have a plan on how to manage it”
13:14:26 bauzas efried: if the operator does that, then any bug related to that should be Wontfix
13:14:47 efried bauzas Well...
13:14:53 efried I agree in principle.
13:14:56 bauzas efried: lemme find where we say that we don't support direct hypervisor calls
13:15:11 efried But that doesn't mean we won't try to accomodate out-of-band partitions when reporting inventory/usage.
13:15:20 openstackgerrit Matt Riedemann proposed openstack/nova master: Log consumer uuid when retrying claims in the scheduler https://review.openstack.org/507705
13:15:33 efried bauzas Sure, but I totally believe you.
13:16:18 openstackgerrit Andrey Volkov proposed openstack/nova master: [WIP] List instances performace optimization https://review.openstack.org/507854
13:17:28 bauzas efried: I didn't found any explanations either in https://docs.openstack.org/nova/latest/contributor/policies.html or https://docs.openstack.org/nova/latest/contributor/project-scope.html
13:17:41 bauzas maybe we should say that
13:17:46 efried cdent Okay, in general, we have very similar issues as described in comment #6 in that bug.
13:18:18 efried When I started looking into converting over to get_inventory, it was going to be a matter of reporting "reserved" amounts based on whatever's going on OOB.
13:18:56 cdent efried: do you have a convenient way of distinguishing between nova managed and not-nova managed stuff?
13:18:56 johnthetubaguy I thought we were heading down the not doing live updates of resource usage?
13:19:04 efried ...and it's not trivial to figure out what's OOB.
13:19:09 efried (I was just about to say :)
13:19:10 cdent jinx!
13:19:35 cdent johnthetubaguy: that seems to work okay for libvirt, but not so great otherwise
13:19:35 efried I can know off the bat to account for my Novalink partition (the node on which the compute service runs) and my Virtual I/O Servers.
13:20:03 efried But any plain ol' worker bees I have to figure out whether they came from Nova.
13:20:17 johnthetubaguy cdent: I guess I am not seeing why, is that vmware bug the example?
13:20:19 openstackgerrit Rodolfo Alonso Hernandez proposed openstack/os-vif master: Migrate from 'ip' commands to 'pyroute2' https://review.openstack.org/484386
13:20:26 cdent johnthetubaguy: at this stage efried, rgerganov and I are just sort of having a chat, not making any decisions
13:20:58 bhagyashris johnthetubaguy, mriedem: Hi,
13:21:01 johnthetubaguy I am just curious where the problem isn't nova driver expected resource usage
13:21:01 efried Simplest would be, every time get_inventory/get_available_resource is called, I ask Nova for the full list of instances it knows about, ask my hypervisor for the instances IT knows about, and do a set-diff.
13:21:20 cdent johnthetubaguy: that’s a specific case of a general issue of “things other than the nova-compute node using the stuff that nova-compute is also using”
13:21:34 efried just so.
13:21:55 johnthetubaguy I think I need a more concrete example
13:22:01 mriedem cdent: i think the vcenter vms have some nova metadata associated with them, so you can tell which ones are nova-managed and which were created oob
13:22:20 efried johnthetubaguy You try to spawn an instance and ask for 3 VCPU. Resource tracker reports you've got 3 VCPU, so the claim passes.
13:22:34 efried johnthetubaguy But in fact, one VCPU is being consumed by a VM that was spawned out of band
13:22:36 rgerganov mriedem, that may work for vcpu and memory but it won't work well for storage
13:22:45 johnthetubaguy its the out of band VMs, I thought we explicitly didn't support that
13:23:03 efried johnthetubaguy Yeah, bauzas said the same, but didn't find where that's documented.
13:23:11 efried Not saying that means it's supported, or that it should be :)
13:23:40 bauzas so, Nova isn't a proxy layer for hypervisors
13:23:49 johnthetubaguy yeah, I could have swarn it was in here, but I don't see it: https://docs.openstack.org/nova/latest/contributor/project-scope.html
13:24:08 bauzas johnthetubaguy: yeah, I verified that
13:24:14 bauzas lemme provide a change for it
13:24:25 johnthetubaguy well, its a bit late I guess
13:25:13 cdent So, even if the statement is “we don’t do that” the problem still holds for shared storage
13:25:21 efried right
13:25:22 mriedem nova doesn't import existing vms on the hypervisor,
13:25:37 cdent where “the problem” is the generic notion of mixed accounting
13:25:37 mriedem but that doesn't mean we don't try to adjust inventory based on things running on the hypervisor host
13:25:39 efried mriedem Which actually makes it more problematic.
13:25:40 efried right.
13:25:43 mriedem that's why we have reserved
13:25:43 johnthetubaguy cdent: yeah, that sounds like a real thing we have to support
13:25:48 efried right
13:26:08 mriedem even libvirt hosts have to account for things like ovs running on the same host
13:26:09 cdent Right, so one of the underlying questions is:
13:26:21 cdent Do we intend/expect that reserved will by dynamically adjuted, frequently
13:26:43 cdent Or in the cases where we want it to be dynamically adjusted we should instead make allocations, via some third party?
13:26:45 efried Right, back to what johnthetubaguy mentioned earlier: "not doing live updates of resource usage"
13:27:06 efried That is, will get_inventory() eventually be a thing that's run only once, rather than on a periodic?
13:27:15 cdent exactly
13:27:17 bhagyashris johnthetubaguy, mriedem: Could you please review patch: https://review.openstack.org/#/c/409644/ ? Addressed all review comments. Thank you :)
13:27:23 mriedem i don't think it will no
13:27:27 bauzas I'm not seeing the reserved bit to be that dynamicv
13:27:31 mriedem the operator can adjust reserved space dynamically
13:27:35 johnthetubaguy cdent: I separate allocations are easier to update than a single reserved value
13:27:42 efried I've heard tell (possibly from jay) that it's a no-no to update the *total* amounts.
13:27:44 mriedem we talked about this in boston
13:27:56 bauzas but yeah, the operator can just set or raise the reserved bit when they want
13:27:58 mriedem about how we don't account for overhead, and operators would have to handle that by toggling reserved values
13:28:23 cdent mriedem: yes, but the new ingredient here is the dynamism
13:28:25 cdent maybe
13:28:27 bauzas mriedem: do you think we should clearly state that in https://docs.openstack.org/nova/latest/contributor/project-scope.html ?
13:28:40 efried ...unless, like, they hot-plug a whole new disk, or jack up the disk size on the SAN or whatever. In that case, we should be able to bump the totals, yes?
13:28:43 bauzas the fact that Nova doesn't support direct hypervisor VMs, but can leave room for them
13:28:49 cdent efried: yes
13:28:56 efried Same actually applies to CPUs
13:29:30 efried Not sure if more primitive hypervisors have this, but there's a dynamic entitlement thing where you can unlock previously-unavailable CPUs on the fly.
13:29:35 johnthetubaguy bhagyashris: there was a -1 review on that patch with no answer when I last looked, so I skipped looking any deeper, did you respond to their questions yet?
13:30:28 johnthetubaguy mriedem: it totally feels like the toggling reserved values ticks the 80% case, with very few surpizes
13:31:02 cdent efried, mriedem, bauzas, johnthetubaguy, rgerganov: it feels like we probably have the tooling to make something workable for this, but as we figure it out it would be good to document some kind of (hate this phrase) best practice,
13:31:32 efried Me, I'm concerned about the penalty of getting the full list of instances every time get_inventory is run.
13:32:35 johnthetubaguy its feels like making the 80% case work really well should be step 1? And some rough best practices for the other bits for now, that could work.
13:33:01 cdent efried: instances from the hypervisor’s point of view, nova’s point of view or placement’s point of view? or some combination? That is, which one worries you most?
13:33:18 johnthetubaguy honestly doing both of those worry me
13:33:46 efried Yeah, I have to do both.
13:33:47 johnthetubaguy I like the system doing nothing when its idle, in an ideal world of course.
13:33:58 bauzas we shouldn't really create a way where people would directly call their VMs if they wish, because that would mean we would silently say we support that
13:34:20 efried Cause the alternative is maintaining some kind of cache and trying to keep it up to date every time a VM gets spawned or deleted (I can get notifications for the OOB ones, and update the in-band ones from spawn/destroy).
13:34:38 bhagyashris johnthetubaguy: Actually I have not respond yet, but Andrey Volkov is given -1 and he suggested that we can fix this issue at schema level (same thing I have proposed in the patch set 3) but after some discussion we decided to fix this in logic and at schema level.
13:34:44 bauzas if the operator doesn't know the dedicated amount their wanna reserve, then either they should disable some compute from the pool, or just set a high value
13:35:18 johnthetubaguy but isn't the easy case much easier here? deployer sets the reserved values, no period updates, worry about special cases later?
13:35:19 bauzas providing a way to dynamically adjust resources in Nova would just mean we create things to support
13:35:23 cdent johnthetubaguy, bauzas: perhaps sad, but true, there are in tree hypervisors where oob vms are fairly common. nova in itself may not need to support that, but the hypervisors will likely continue doing it
13:35:25 bhagyashris s/and at schema level/and not at schema level
13:35:50 bauzas johnthetubaguy: that's my point
13:36:08 dansmith mriedem: have you seen the stuff I added to the etherpad last night?
13:36:22 bauzas johnthetubaguy: if operators really want a space for out-of-band VMs, then they have the reserved bit they can set and try to modify when they feel necessary
13:36:43 bhagyashris johnthetubaguy: i will respond to that question.
13:36:54 bauzas but trying to provide some mechanism for having a dynamic reserved bit seems to me something we shouldn't do
13:36:54 johnthetubaguy cdent: but even those drivers, there are folks who run dedicated clusters just for OpenStack
13:37:41 bauzas and honestly, wouldn't it be better to just disable a host and play with it, if that's really needed ?

Earlier   Later