| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-05-11 | |||
| 11:02:21 | gibi | lyarwood: for me either the size of the cache should be configurable (and then we can account for that in reserved_host_disk_mb) or nova needs to report cached disk usage as used in placement | |
| 11:02:52 | lyarwood | gibi: yeah I'd say the latter, I honestly thought we did already. | |
| 11:04:25 | gibi | lyarwood: this is a bug in pike for one of our customers so I also have to think about a backportable solution | |
| 11:04:58 | gibi | lyarwood: do you think that determining the size of the image cache is easy? | |
| 11:05:18 | gibi | is it just some file system calls in the _base dir, isn't it? | |
| 11:05:48 | lyarwood | gibi: yes for file based backends | |
| 11:06:27 | gibi | I have close to zero knowledge on non file based backends behavior | |
| 11:07:13 | lyarwood | gibi: we don't cache in rbd iirc | |
| 11:07:33 | gibi | is there a way to turn of the cache? | |
| 11:07:53 | gibi | for file based backend? (that would be a workaround in my downstream issue) | |
| 11:09:48 | lyarwood | gibi: I don't think so, you can disable the manager but I think we still cache things at creation time | |
| 11:10:00 | lyarwood | gibi: the manager just doesn't run to clean things up | |
| 11:10:20 | gibi | lyarwood: thanks | |
| 11:57:19 | openstackgerrit | Sasha Andonov proposed openstack/nova master: rbd_utils: increase _destroy_volume timeout https://review.opendev.org/705764 | |
| 12:07:09 | nightmare_unreal | how can I determine default api version used by osc CLI ? suppose if we don't specify --os-compute-api-version , which version will it take and how it's determined ? | |
| 12:10:14 | lyarwood | nightmare_unreal: https://docs.openstack.org/api-guide/compute/microversions.html#version-discovery - I think there's also an osc command for that | |
| 12:10:36 | nightmare_unreal | Thanks :) | |
| 12:52:10 | openstackgerrit | Stephen Finucane proposed openstack/nova master: docs: Note the ``hw_numa_nodes`` image property https://review.opendev.org/683849 | |
| 13:34:25 | openstackgerrit | Stephen Finucane proposed openstack/nova master: SR-IOV passthrough: Check PF only if VF is enabled https://review.opendev.org/476642 | |
| 13:36:19 | jsuchome | dansmith: Hi, I've updated related patchets about that direct rbd download (including the spec) could you give it another look? | |
| 13:46:17 | openstackgerrit | Stephen Finucane proposed openstack/nova master: Poison netifaces.interfaces() in tests https://review.opendev.org/671773 | |
| 13:46:51 | dansmith | jsuchome: yep will queue for today | |
| 13:51:52 | openstackgerrit | Stephen Finucane proposed openstack/nova master: Raise if flavor and image disagree on hide_hypervisor_id https://review.opendev.org/663365 | |
| 14:17:16 | openstackgerrit | Lee Yarwood proposed openstack/nova stable/train: Revert "nova shared storage: rbd is always shared storage" https://review.opendev.org/726861 | |
| 14:17:35 | openstackgerrit | Lee Yarwood proposed openstack/nova stable/stein: Revert "nova shared storage: rbd is always shared storage" https://review.opendev.org/726862 | |
| 14:17:50 | openstackgerrit | Lee Yarwood proposed openstack/nova stable/rocky: Revert "nova shared storage: rbd is always shared storage" https://review.opendev.org/726863 | |
| 14:18:03 | openstackgerrit | Lee Yarwood proposed openstack/nova stable/queens: Revert "nova shared storage: rbd is always shared storage" https://review.opendev.org/726864 | |
| 14:23:41 | sean-k-mooney | by the way are we planning to backport https://review.opendev.org/#/c/663365/ upstream | |
| 14:24:02 | sean-k-mooney | its the fix for bug@ #1831723 | |
| 14:24:18 | sean-k-mooney | bug: #1831723 | |
| 14:24:18 | openstack | bug 1831723 in OpenStack Compute (nova) "The flavor hide_hypervisor_id value can be overridden by the image img_hide_hypervisor_id" [Undecided,In progress] https://launchpad.net/bugs/1831723 - Assigned to Stephen Finucane (stephenfinucane) | |
| 14:24:50 | sean-k-mooney | downstream i think we would want to backport that as im sure customer will hit it at some point | |
| 14:25:10 | sean-k-mooney | but it might be nice to backport upstream too but im not sure its allowed | |
| 14:25:24 | sean-k-mooney | lyarwood: ^ any toughts on the topic | |
| 14:26:07 | sean-k-mooney | we are just adding a namespaced version of an existing unnamsepaced extra_spec | |
| 14:26:29 | sean-k-mooney | downstream that is not conisderd an api change since extra_specs are not part of the api | |
| 14:26:35 | lyarwood | sean-k-mooney: reading | |
| 14:27:34 | openstackgerrit | James Page proposed openstack/nova stable/queens: hardware: fix memory check usage for small/large pages https://review.opendev.org/726867 | |
| 14:28:58 | sean-k-mooney | lyarwood: actully that is not the patch i ment to link | |
| 14:29:34 | lyarwood | sean-k-mooney: right you had me slightly confused tbh | |
| 14:29:50 | sean-k-mooney | that one we might also want to backport but one sec | |
| 14:29:58 | lyarwood | yeah that's looks valid to backport | |
| 14:30:01 | lyarwood | that* | |
| 14:30:04 | sean-k-mooney | https://review.opendev.org/#/c/722187/ | |
| 14:30:08 | sean-k-mooney | that is the one i ment | |
| 14:30:27 | sean-k-mooney | for https://bugs.launchpad.net/nova/+bug/1841932 | |
| 14:30:27 | openstack | Launchpad bug 1841932 in OpenStack Compute (nova) "hide_hypervisor_id extra_specs in nova flavor cannot pass AggregateInstanceExtraSpecsFilter" [Low,In progress] - Assigned to Stephen Finucane (stephenfinucane) | |
| 14:33:03 | lyarwood | sean-k-mooney: so we deprecate hide_hypervisor_id in that change but still provide backward compatability so at first glance I think we can backport this upstream? | |
| 14:33:36 | sean-k-mooney | lyarwood: correct deprecated but still supported | |
| 14:33:49 | sean-k-mooney | we would have to drop the validation changes | |
| 14:33:53 | sean-k-mooney | form the patch | |
| 14:34:07 | sean-k-mooney | but the rest of it i think would be fine | |
| 14:34:47 | lyarwood | ah right that only just landed | |
| 14:34:50 | openstackgerrit | James Page proposed openstack/nova stable/queens: hardware: fix memory check usage for small/large pages https://review.opendev.org/726867 | |
| 14:34:51 | openstackgerrit | James Page proposed openstack/nova stable/queens: Fix overcommit for NUMA-based instances https://review.opendev.org/726868 | |
| 14:35:27 | sean-k-mooney | lyarwood: yep but that should be fine and easy to call out on the intiall backport patch | |
| 14:35:32 | sean-k-mooney | its pretty self contained | |
| 15:01:03 | jsuchome | dansmith: thanks, I'll address the changes in spec | |
| 15:12:29 | gibi | lyarwood: reported a bug for the image cache issue https://bugs.launchpad.net/nova/+bug/1878024 | |
| 15:12:29 | openstack | Launchpad bug 1878024 in OpenStack Compute (nova) "disk usage of the nova image cache is not counted as used disk space" [Undecided,New] | |
| 15:12:58 | gibi | lyarwood: could you please check if it make sense what I wrote there? | |
| 15:13:30 | gibi | dansmith: you worked with the image cache recently so you might be interested ^^ | |
| 15:13:37 | dansmith | yeah, reading now | |
| 15:13:52 | gibi | thanks | |
| 15:13:57 | dansmith | if the op has a separate mount for the cache, this wouldn't be a problem | |
| 15:14:09 | dansmith | so it's not going to affect everyone the sae | |
| 15:14:10 | dansmith | *same | |
| 15:14:30 | gibi | yeah, what the downstream customer has is a simple disk partition for the nova instances_path | |
| 15:14:41 | gibi | and the cache is under the instances_path | |
| 15:14:43 | dansmith | right, which I'm sure is common | |
| 15:14:57 | lyarwood | gibi: ack thanks just on calls for a while but will look once I'm off | |
| 15:15:04 | gibi | lyarwood: thanks | |
| 15:15:39 | dansmith | the problem with something like A is that when you're scheduling, | |
| 15:15:58 | dansmith | you don't know whether or not the image is already on the remote system, so you don't know whether to look for hosts with 2*$size disk space or not | |
| 15:16:39 | gibi | dansmith: true | |
| 15:16:43 | dansmith | B doesn't really work either because you can't assume you can purge your way out of the cache limit | |
| 15:17:02 | dansmith | if you boot a hundred instances from different images, you can't prevent the image cache from going over the desired size, | |
| 15:17:19 | dansmith | unless you refuse to boot instances there, which people will complain about because there is plenty of disk space and not understand | |
| 15:18:11 | gibi | could nova-compute periodically update a separate DISK_GB allocation in placement based on the actual size of the image cache | |
| 15:18:17 | gibi | ? | |
| 15:18:52 | dansmith | not periodically, but synchronously with the decision to cache an image (either during boot or otherwise) | |
| 15:18:57 | dansmith | otherwise you have a race | |
| 15:20:02 | dansmith | I'd have to think about that, but we'd need to only do that if the images and cache are on the same filesystem, otherwise we'd count against the wrong total | |
| 15:20:13 | gibi | it will be racey anyhow due to what you said about the problem of requesting allocation for the cache during scheduling | |
| 15:21:07 | dansmith | I mean racing for disk space, which could go badly if you lose, not just racing for image boot | |
| 15:21:30 | dansmith | but yes, the scheduler is never going to know whether or not to pick a host based on cache availability, so you always have that | |
| 15:22:24 | dansmith | we could go totally crazy and create an allocation for each image, by image uuid and after selecting a host, the scheduler could check to see if there was an allocation for that image against the host's provider to decide if it thinks it will fit :) | |
| 15:22:43 | dansmith | that has some nice benefits, but it's a little crazy and there's still plenty of room for racing of course, | |
| 15:23:03 | dansmith | and plenty of room for exhausting all the candidates in a small query set, leading to non-ideal looping of retries | |
| 15:23:39 | dansmith | plus we'd have to have a separate provider for the cache disk if they're separate | |
| 15:27:07 | gibi | complicated indeed | |
| 15:27:41 | gibi | I will pass the workaround of having the cache on a different partition to downstream | |
| 15:27:56 | gibi | at least that is something that the downstream project can do | |
| 15:28:24 | aarents | gibi: dansmith we have this issue, painfull one | |
| 15:28:48 | dansmith | gibi: ack | |
| 15:29:02 | aarents | gibi yep in some case we put cahe in another file system to get ride of this | |
| 15:29:22 | gibi | dansmith: the old DiskFilter had the disk_available_least info to prevent overallocation but we removed the DiskFilter | |
| 15:30:27 | dansmith | yeah, fair point | |
| 15:30:56 | gibi | we might want to re-introduce something like diks_available_least as a filter? or a pre-filter with placement support? | |
| 15:31:40 | dansmith | I think diskfilter had plenty of other problems, like the other way where the filter behavior conflicted with the hypervisors listing, which definitely causes support cases | |