| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2021-01-13 | |||
| 17:32:47 | melwitt | thank you stephenfin | |
| 17:33:29 | markguz_ | why would the lock take 25mins? | |
| 17:34:15 | sean-k-mooney | the comput service is likely bussy starting up | |
| 17:34:24 | sean-k-mooney | you mentioned this is only after the iniall start right | |
| 17:34:41 | sean-k-mooney | or is this for each spawn | |
| 17:34:59 | lyarwood | just grep for the compute_resources lock and see what was holding it before? | |
| 17:35:35 | sean-k-mooney | ya that too | |
| 17:35:45 | sean-k-mooney | you could see how many other instance got the lock in that interval | |
| 17:36:22 | markguz_ | sean-k-mooney: it's every spawn. usually 10mins, sometimes longer and sometimes shorter | |
| 17:36:55 | lyarwood | does the resource tracker make external API calls in the Ironic driver? | |
| 17:38:46 | sean-k-mooney | the RT is shared so i dont think so | |
| 17:39:47 | lyarwood | yeah sorry I mean the code that refreshes it within the driver | |
| 17:40:11 | sean-k-mooney | im pretty sure this is the lock in question https://opendev.org/openstack/nova/src/branch/stable/rocky/nova/compute/manager.py#L2221-L2222 | |
| 17:41:13 | melwitt | markguz_: I think you might be hitting https://bugs.launchpad.net/nova/+bug/1864122 | |
| 17:41:15 | openstack | Launchpad bug 1864122 in OpenStack Compute (nova) "Instances (bare metal) queue for 30-60 seconds when managing a large amount of Ironic nodes" [Medium,Fix released] - Assigned to Jason Anderson (jasonandersonatuchicago) | |
| 17:41:50 | sean-k-mooney | yep that was what grabed the lock https://opendev.org/openstack/nova/src/branch/stable/rocky/nova/compute/resource_tracker.py#L159-L160 | |
| 17:42:13 | markguz_ | i have +/- 240 nodes | |
| 17:42:24 | markguz_ | is that a large amount? | |
| 17:42:37 | sean-k-mooney | melwitt: yep a race on the lock with the update periodic task seams likely | |
| 17:42:41 | melwitt | if you read the bug it says can be seen around > 100 nodes | |
| 17:43:12 | sean-k-mooney | markguz_: how many ironic compute services do you have | |
| 17:43:15 | melwitt | that fix is available in ussuri and onward, it was not backported because it requires a newer version of oslo.concurrency | |
| 17:43:24 | markguz_ | sean-k-mooney: 1 | |
| 17:43:29 | sean-k-mooney | i belive the periodic will only update the resouce usage for the nodes that are assgined to it | |
| 17:44:10 | sean-k-mooney | so i think you can scale it by deploying more compute service instances TheJulia is that correct? | |
| 17:45:03 | sean-k-mooney | markguz_: if you have 3 contolers i would suggest running an ironic nova compute service instance on each assuming that makes sense to TheJulia or others | |
| 17:45:07 | TheJulia | sean-k-mooney: yes, you can, you should just be able to run multiple instances | |
| 17:45:19 | TheJulia | markguz_: ^^^ instances of nova-compute configured for ironic | |
| 17:45:33 | stephenfin | melwitt: comments left on the bug report too | |
| 17:46:09 | sean-k-mooney | markguz_: the other thing you could do is reduce the interval of the periodic | |
| 17:46:21 | melwitt | hm, I thought you needed to configure node partitioning to do that | |
| 17:46:23 | sean-k-mooney | we fixed it by chanigin the type of lock we use | |
| 17:46:46 | melwitt | "conductor groups" | |
| 17:46:52 | TheJulia | melwitt: only to force specific grouping/allocation into specific grouping | |
| 17:46:56 | sean-k-mooney | that on the ironic side i think | |
| 17:47:01 | melwitt | it's not | |
| 17:47:05 | TheJulia | its on both sides | |
| 17:47:12 | sean-k-mooney | ah ok | |
| 17:47:19 | markguz_ | peridoc_task_interval is set to 240 | |
| 17:47:19 | melwitt | well, it might be but you have to do it on the nova side too | |
| 17:47:24 | TheJulia | otherwise it runs a hash ring based upon the node list | |
| 17:47:31 | TheJulia | and the group is just a key in the hash ring | |
| 17:48:01 | sean-k-mooney | we improved this in nova by using oslos fair locks | |
| 17:48:02 | TheJulia | the nova side name is a little different because naming_is_fun^TM | |
| 17:48:02 | sean-k-mooney | https://review.opendev.org/c/openstack/nova/+/711528/2 | |
| 17:48:23 | markguz_ | so we use this in a lab env and when we spin up baremetal we need to spin up a spcific node as they are connected to specific hardware that is being tested | |
| 17:48:39 | sean-k-mooney | but that was only done in ussuri | |
| 17:48:48 | TheJulia | sean-k-mooney: ohhhhh neat | |
| 17:49:02 | sean-k-mooney | so we would have to backport it unfortunetlly im not sure oslo has the required support in rocky let me check | |
| 17:49:13 | melwitt | ok conductor groups are not available until stein anyways | |
| 17:49:30 | sean-k-mooney | we would need oslo.concurrancy 3.29.0 to backport it | |
| 17:49:39 | markguz_ | if i run multiple computes i'm guessing that i will need to change how i call a instance. right now i use the "avail_zone:compute_host:bm_uuid" trick | |
| 17:49:40 | melwitt | yeah, I said all of that earlier | |
| 17:49:55 | sean-k-mooney | stable rocky is oslo.concurrency===3.27.0 | |
| 17:50:07 | TheJulia | markguz_: yeah, :\ | |
| 17:50:14 | melwitt | yes, the patch that added fair locks bumped the oslo.concurrency version | |
| 17:50:22 | sean-k-mooney | so it can go back to stien | |
| 17:50:32 | sean-k-mooney | but not rocky | |
| 17:50:40 | melwitt | so it wasn't bumped until ussuri | |
| 17:51:18 | sean-k-mooney | markguz_: ya you would need to know which host has it | |
| 17:51:33 | markguz_ | what's ironic is (pun intended) is that i was upgrading with the intention of getting to ussuri but when ironic failed at the rocky step i didn't want to compound the problem by continuing to upgrade | |
| 17:52:37 | melwitt | huh yeah actually it could be backported to stein because the upper constraint is 3.29.1 for whatever reason | |
| 17:52:56 | melwitt | I did not expect that | |
| 17:53:02 | markguz_ | assuming the bug is the problem going to ussuri will fix it? but that will break a lot of our automation due to the way were calling the nodes | |
| 17:53:25 | markguz_ | i mean it's not the end of the world, but ugh.. more work :-( | |
| 17:53:58 | markguz_ | at least i finally have a better idea of what's wrong at least. i seriously was losing the will to live over this ;-) | |
| 17:54:19 | melwitt | yeah. you could try to haxx and apply the patch to see if it helps. you just need oslo.concurrency >= 3.29.0 | |
| 17:54:45 | melwitt | (so you know for sure whether you're hitting that bug) | |
| 17:55:15 | markguz_ | melwitt: does the patch need to go on the scheduler or the compute node? or both? | |
| 17:55:25 | melwitt | markguz_: compute node | |
| 17:56:15 | markguz_ | melwitt: then i can probably crowbar that in | |
| 17:58:12 | melwitt | bleh, there's merge conflicts but it's really just adding fair=True to all the @utils.synchronized(COMPUTE_RESOURCE_SEMAPHORE, fair=True) | |
| 17:59:03 | markguz_ | ok. i'll give it try and see what happens. | |
| 18:00:03 | markguz_ | will i need to upgrade the other oslo. components or just concurrency? | |
| 18:00:40 | melwitt | just concurrency | |
| 18:02:03 | openstackgerrit | melanie witt proposed openstack/nova stable/train: Use fair locks in resource tracker https://review.opendev.org/c/openstack/nova/+/770585 | |
| 18:04:32 | sean-k-mooney | we could proably implenet a version of the patch for rocky too | |
| 18:04:40 | sean-k-mooney | that just did not use the fair lock form oslo | |
| 18:10:23 | sean-k-mooney | markguz_: this was the implemenation fo the fair lock https://github.com/openstack/oslo.concurrency/commit/2b55da68ae45ff45cba68672cdbc24342cf115f6 | |
| 18:13:32 | sean-k-mooney | markguz_: if you wanted to backport the upstream patch and then backport the implemantion of the fair lock into nova we could evaulate that or at least the stable team could | |
| 18:13:53 | sean-k-mooney | its just using https://fasteners.readthedocs.io/en/latest/api/lock.html#fasteners.lock.ReaderWriterLock | |
| 18:14:03 | sean-k-mooney | to actuly provide the fifo behavior | |
| 18:15:50 | sean-k-mooney | the version of fasteners on stable rocky has the required functionality | |
| 18:16:45 | sean-k-mooney | that said i think you can just bump the oslo.concurrancy version locally and locally apply the nova patch and it should run fine | |
| 18:19:52 | openstackgerrit | melanie witt proposed openstack/nova stable/stein: Use fair locks in resource tracker https://review.opendev.org/c/openstack/nova/+/770657 | |
| 18:20:40 | sean-k-mooney | stephenfin: by the way as far as i am aware we never us the cpu_toplopgy filed in the numa cell object to generate teh xml at all | |
| 18:20:47 | openstackgerrit | melanie witt proposed openstack/nova stable/stein: Use fair locks in resource tracker https://review.opendev.org/c/openstack/nova/+/770657 | |
| 18:21:58 | sean-k-mooney | the numa toplogy of the guest or host should have no impact on the cpu toplogy of the guest period | |
| 18:22:28 | sean-k-mooney | any other behviaor is inconsistent with the intended behviaor as discibed by the specs | |
| 18:58:19 | openstackgerrit | Merged openstack/nova stable/victoria: Omit resource inventories from placement update if zero https://review.opendev.org/c/openstack/nova/+/766177 | |
| 19:14:38 | markguz_ | sean-k-mooney: i think it would be simpler for me to just upgrade to ussuri | |
| 19:14:58 | sean-k-mooney | if that is an option yes | |
| 19:15:37 | openstackgerrit | Merged openstack/nova stable/victoria: Add upgrade check about old computes https://review.opendev.org/c/openstack/nova/+/761924 | |
| 19:15:43 | markguz_ | fortunately for me this is an internal deployment that is not used by paying customers so i have some degree of flexibility on it's availability | |
| 19:16:23 | sean-k-mooney | what do you use to deploy/manage it | |
| 19:21:11 | markguz_ | sean-k-mooney: originally i deployed kilo with rdo packstack. since then it's become a bit of a bit of hodgepodge of manual installs. I mostly use ansible to keep things up to date | |
| 19:21:48 | sean-k-mooney | ah i see | |
| 19:22:01 | sean-k-mooney | packstack has more or less been unsupported for a few years now | |
| 19:22:27 | sean-k-mooney | i think it still technicaly exists but redhat stop supporting in with our product in queens i think | |
| 19:22:47 | markguz_ | yeah. i generally just install from the package manager. and have some ansible plays that configure compute nodes etc etc. | |