| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-27 | |||
| 13:36:54 | johnthetubaguy | cdent: but even those drivers, there are folks who run dedicated clusters just for OpenStack | |
| 13:37:41 | bauzas | and honestly, wouldn't it be better to just disable a host and play with it, if that's really needed ? | |
| 13:37:44 | cdent | johnthetubaguy: yes, totally | |
| 13:37:47 | johnthetubaguy | cdent: that was me arguing for making the sync optional, for when you know Nova owns the world | |
| 13:37:58 | rgerganov | johnthetubaguy, most users runs dedicated clusters but also used shared storage | |
| 13:38:07 | dansmith | johnthetubaguy: are you talking about the periodic that will try to delete unknown vms? | |
| 13:38:09 | mriedem | dansmith: sure did | |
| 13:38:14 | johnthetubaguy | bhagyashris: thanks, responding in the review will help wehn I get back to that | |
| 13:38:21 | mriedem | dansmith: nope, talking about update_available_resource | |
| 13:38:41 | johnthetubaguy | dansmith: well, I was thinking about that early, with all the logging by default, but there I was meaning the resource tracker updater | |
| 13:38:41 | mriedem | and using get_inventory to report reserved space for other things running on the hypervisor outside of nova's control | |
| 13:39:10 | dansmith | mriedem: eff that. there's a reserved count for that kind of stuff. | |
| 13:39:17 | bhagyashris | johnthetubaguy: Thank you :) | |
| 13:39:34 | mriedem | dansmith: meaning the config option? | |
| 13:39:48 | dansmith | meaning counting other things and subtracting it ourselves | |
| 13:40:02 | mriedem | i think the config option is what's used to report reserved inventory in the RT today | |
| 13:40:13 | dansmith | yes | |
| 13:40:33 | dansmith | well, actually, I dunno that it is today, but it should be, and that should be your outlet to carve out space | |
| 13:41:12 | dansmith | it is | |
| 13:41:17 | dansmith | https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L102-L125 | |
| 13:41:39 | dansmith | make those hot-reloadable if you want, but that's the path, IMHO | |
| 13:42:58 | efried | dansmith That guy still gets run on a periodic basis, right? I.e. dynamic tweaking of reserved amounts is supported (at least by the driver - not asking if those conf opts are dynamic) | |
| 13:43:31 | johnthetubaguy | we should make SIG_HUP trigger a resource refresh? | |
| 13:43:46 | dansmith | johnthetubaguy: HUP already triggers a config re-read | |
| 13:43:54 | johnthetubaguy | I was meaning a resource tracker update | |
| 13:44:03 | dansmith | johnthetubaguy: so if they were reloadable you'd get a new value the next time inventory runs | |
| 13:44:07 | johnthetubaguy | so we don't have to keep the periodic thing, just do it on events | |
| 13:44:30 | dansmith | I'm not sure there's any reason not to do it perioically | |
| 13:44:53 | johnthetubaguy | system load, all that lists of instances, etc | |
| 13:45:28 | dansmith | it doesn't need to anymore | |
| 13:46:00 | dansmith | in pike, once there aren't any ocata computes around we're not doing healing for instances other than deleted ones | |
| 13:46:31 | johnthetubaguy | I guess I am confused | |
| 13:47:24 | johnthetubaguy | I thought we were left with no periodic updates in queens | |
| 13:47:31 | johnthetubaguy | to inventory, etc | |
| 13:47:44 | johnthetubaguy | although... ironic needs those | |
| 13:47:52 | johnthetubaguy | dang it | |
| 13:48:19 | dansmith | I'm saying, we can remove a bunch of the RT stuff for healing instances now I think, | |
| 13:48:34 | johnthetubaguy | ah, OK, so maybe I am agreeing with you | |
| 13:48:35 | dansmith | which will make the periodic mostly an inventory mechanism, which doesn't need to be listing instances and such | |
| 13:49:32 | efried | The periodic will use driver.get_inventory() to update inventories? | |
| 13:49:57 | efried | (still) | |
| 13:50:04 | dansmith | or get_available_resource() if it doesn't have that | |
| 13:50:09 | efried | Cool. | |
| 13:50:10 | dansmith | or whatever that method is called | |
| 13:50:49 | efried | So I think the point is that, in order to do dynamic adjustment of 'reserved' amounts properly, get_inventory itself may need to be listing instances and such. | |
| 13:50:59 | dansmith | no | |
| 13:51:07 | dansmith | inventory has nothing to do with instances | |
| 13:51:42 | rgerganov | dansmith, there could be no way to calculate reserve without listing instances | |
| 13:52:06 | dansmith | reserved is a conf value | |
| 13:52:44 | bauzas | that's my whole point | |
| 13:52:46 | dansmith | you all were talking about other vms not owned by nova on a hypervisor right? and something external adding that value to reserved so nova doesn't try to use it, right? | |
| 13:52:53 | bauzas | reserved is set by the operator | |
| 13:53:08 | bauzas | so, if the operator wants to raise that value, fair enough | |
| 13:53:18 | dansmith | in that case, some other thing could be updating CONF.reserved_whatever and HUPing nova so that the next time the periodic runs, it just stabs the new scalar value into inventory.. no counting required. | |
| 13:53:19 | bauzas | but I don't see why Nova should care of thaty | |
| 13:53:24 | dansmith | bauzas: most definitely | |
| 13:53:26 | bauzas | dansmith: +1 | |
| 13:55:10 | mriedem | bauzas: in case you didn't see this https://bugs.launchpad.net/nova/+bug/1719730 | |
| 13:55:12 | openstack | Launchpad bug 1719730 in OpenStack Compute (nova) pike "Reschedule after the late affinity check fails with "'NoneType' object is not iterable"" [High,Confirmed] | |
| 13:55:19 | mriedem | bauzas: melwitt might be working on a fix, i'm not sure | |
| 13:55:19 | efried | As user experiences go, "Extend SAN disk, edit conf file, HUP n-cpu" ain't as friendly as "Extend SAN disk". | |
| 13:55:24 | mriedem | bauzas: but it's a regression in pike | |
| 13:55:32 | bauzas | mriedem: yup, I just provided my thoughts litterally 5 mins ago :p | |
| 13:55:53 | bauzas | mriedem: tl;dr I don't understand how we can get a ReqSpec group info that isn't accurate | |
| 13:55:54 | rgerganov | efried +1 | |
| 13:56:12 | dansmith | efried: nova wouldn't be reporting the inventory for a SAN disk, so no problem there :) | |
| 13:56:30 | dansmith | efried: and anything that an agent is surveying would be updated automatically | |
| 13:56:52 | dansmith | efried: if you hotplug some memory into your compute node, then nova-compute would automatically report that on the next periodic | |
| 13:57:19 | dansmith | efried: what you are talking about is updating the reserved value, which is static, operator-configured or operator-scripted | |
| 13:57:36 | mriedem | bauzas: provided your thoughts where? don't see anything in the bug or the linked pike change | |
| 13:57:39 | dansmith | efried: the HUP is so we notice the change to the config file.. certainly you don't want us to be live reloading the config file whenever we feel like it | |
| 13:57:54 | bauzas | mriedem: oh snap, duplicate bug | |
| 13:58:05 | bauzas | mriedem: https://bugs.launchpad.net/nova/+bug/1719859 | |
| 13:58:06 | openstack | Launchpad bug 1719859 in OpenStack Compute (nova) "Resize failure due to instance group being None in request spec" [Undecided,New] | |
| 13:58:30 | efried | dansmith Right, extending SAN was a bad example. | |
| 13:58:44 | bauzas | mriedem: mmm, nevermind, different issue, it seems | |
| 13:59:34 | efried | Not being super well-versed on how OVS comes into play for libvirt, can we apply the same principle to networking resources? | |
| 13:59:42 | mriedem | bauzas: ok, mel's is pretty straight forward | |
| 13:59:49 | mriedem | we used to set group_members in the filter_properties, | |
| 13:59:56 | mriedem | your change in pike stopped using filter_properties, | |
| 13:59:58 | efried | I.e. similar UX statement: "Do OVS thingy, edit conf file, HUP n-cpu" not as friendly as "Do OVS thingy" | |
| 14:00:02 | mriedem | but forgot to include the group_members stuff | |
| 14:00:12 | mriedem | she already said that the 1 line change fixes the issue | |
| 14:00:20 | mriedem | i just asked that a functional regression test get created | |
| 14:00:24 | bauzas | mriedem: okay, I'll wait for melwitt's patch | |
| 14:00:29 | dansmith | efried: no, nova-compute isn't reporting any network resources | |
| 14:00:53 | bauzas | mriedem: thanks for the catch | |
| 14:01:00 | mriedem | bauzas: which reminds me, did you cleanup that num_instances one? | |
| 14:01:06 | bauzas | mriedem: I got some late pings yesterday | |
| 14:01:08 | mriedem | melwitt found the problem, i just helped debug | |
| 14:01:22 | mriedem | https://review.openstack.org/#/c/506093/ | |
| 14:01:33 | dansmith | efried: I think the block you're stumbling over is that nova does not count things it does not count. Therefore, it does not dynamically update inventory for things it does not count. If you want to override the inventory for a thing it _does_ count to account for things it does not count, then you put that in config an HUP it to notice. | |
| 14:01:42 | bauzas | mriedem: yeah, just needs rebase | |
| 14:01:46 | mriedem | i need to start making an etherpad of bug fixes we need in pike... | |
| 14:02:07 | bauzas | mriedem: I also have a couple of bugfixes that are on hold in my queue | |
| 14:02:17 | bauzas | like the fact we don't destroy the Spec objects | |
| 14:02:42 | bauzas | mriedem: an etherpad seems a good idea to me | |
| 14:02:54 | bauzas | that would also help me focusing on rebasing such bugs | |
| 14:03:10 | mriedem | https://etherpad.openstack.org/p/nova-pike-bug-fix-backports | |
| 14:04:01 | bauzas | ack | |