Earlier  
Posted Nick Remark
#openstack-nova - 2022-09-05
09:20:31 sean-k-mooney[m] or have any other sideffect we would not want
09:23:38 gibi storing row is OK the doc said dict hence my statement that the caches is broken
09:23:55 sean-k-mooney[m] ack
09:24:34 gibi we need to mix Row and namedtuple as I can only create namedtuple here https://github.com/openstack/placement/blob/13bbdba06da19f85c05a2a9e1fbdb9d1813c3b47/placement/attribute_cache.py#L157-L168
09:24:40 gibi but they are compatible
09:24:43 gibi so I only leave a note
09:24:55 gibi when I convert that to namedtuple
09:39:38 sean-k-mooney[m] if you use a named tuple on line 155 as well in refersh_from_db i thik we dont need to mix types but cool ill review when you push
09:47:01 songwenping_ sean-k-mooney[m],gibi: hi, nova-scheduler get allocation_candidates return 504 gateway timeout when create vm with 8gpus requests on our client's env, there are 13 gpu compute, and every compute has 8gpus.
09:47:02 opendevreview Balazs Gibizer proposed openstack/placement master: Make us compatible with oslo.db 12.1.0 https://review.opendev.org/c/openstack/placement/+/855862
09:47:07 gibi sean-k-mooney[m]: ^^
09:47:12 gibi stephenfin: ^^
09:50:49 gibi sean-k-mooney[m]: also here is my stab at the nova only fair lock fix https://review.opendev.org/c/openstack/nova/+/855717 there is the oslo version of the fix https://review.opendev.org/c/openstack/oslo.concurrency/+/855714
09:51:01 songwenping_ i imitate to insert some test datas on my devstack env, and the result is same, perhaps the api of allocation_candidates/limit... need to optimize.
09:51:27 gibi but I have to go back and think about the unit tests as it seems they are unstable
09:53:13 opendevreview Amit Uniyal proposed openstack/nova master: Adds check for VM snapshot fail while quiesce https://review.opendev.org/c/openstack/nova/+/852171
09:54:22 sean-k-mooney[m] songwenping_: this is physical gpu passthoug?
09:54:29 sean-k-mooney[m] not vGPU correct
09:54:30 songwenping_ yes
09:54:41 songwenping_ pgpu passthough
09:55:05 sean-k-mooney[m] those are not tracked in placment then
09:55:18 auniyal_ Hi
09:55:23 auniyal_ please review these
09:55:25 auniyal_ https://review.opendev.org/c/openstack/nova/+/852171
09:55:25 auniyal_ https://review.opendev.org/c/openstack/nova/+/854499
09:55:25 auniyal_ backporting
09:55:25 auniyal_ https://review.opendev.org/c/openstack/nova/+/854979
09:55:25 auniyal_ https://review.opendev.org/c/openstack/nova/+/854980
09:56:17 sean-k-mooney[m] songwenping_: on master we now can track pci devics in placment but in any other release pci devices are not tracked in placment
09:56:34 songwenping_ sean-k-mooney[m]: we use cyborg to manage pgpu, the request url is :curl -g -i -X GET "http://10.7.20.73/placement/allocation_candidates?limit=1000&group_policy=none&required1=CUSTOM_GPU_NVIDIA%2CCUSTOM_GPU_PRODUCT_ID_1DB6&required2=CUSTOM_GPU_NVIDIA%2CCUSTOM_GPU_PRODUCT_ID_1DB6&required3=CUSTOM_GPU_NVIDIA%2CCUSTOM_GPU_PRODUCT_ID_1DB6&required4=CUSTOM_GPU_NVIDIA%2CCUSTOM_GPU_PRODUCT_ID_1DB6&required5=CUSTOM_GPU_NVIDIA%2CCUSTOM_GPU_PRODUCT
09:56:34 songwenping_ _ID_1DB6&required6=CUSTOM_GPU_NVIDIA%2CCUSTOM_GPU_PRODUCT_ID_1DB6&resources=MEMORY_MB%3A32%2CVCPU%3A1&resources1=PGPU%3A1&resources2=PGPU%3A1&resources3=PGPU%3A1&resources4=PGPU%3A1&resources5=PGPU%3A1&resources6=PGPU%3A1" -H "Accept: application/json" -H "OpenStack-API-Version: placement 1.29" -H "User-Agent: openstacksdk/0.99.0 keystoneauth1/4.6.0 python-requests/2.27.1 CPython/3.8.10" -H "X-Auth-Token: gAAAAABjFbkj8nz24Q7A6J0qmjpdZHfWM
09:56:35 songwenping_ vZidJOT9iYhV2MQCngcYQHhSQmjsGJofkYoT087tAISpf3IniDGwPTHXz_-8x-1nF60WavSYFgEd-5l3_ENrGumaHuU1yfhMJqZu06IR4SXacjA1g6ImSSEfLbfQ9zrPouB0roFokHPmPy3-UpnFZE"
09:56:46 gibi sean-k-mooney[m], sean-k-mooney[m]: is it via cyborg? because then it might be tracked in placement
09:56:53 sean-k-mooney[m] ok
09:58:05 sean-k-mooney[m] in that case perhaps they are hitting the compintorial explosion we were worried about with tracking VFs directly
09:58:11 gibi probably there too many possible candidates
09:58:19 sean-k-mooney[m] there are 13 chose 8 combinations
09:58:39 sean-k-mooney[m] that 1287
09:58:46 gibi that is not that much
09:58:55 gibi but if there is 1000 computes
09:58:59 sean-k-mooney[m] 1287*1000
09:58:59 gibi or just 100
09:59:07 gibi then that is sizeable
09:59:16 sean-k-mooney[m] ya it will grow quickly
10:00:03 sean-k-mooney[m] sorry no
10:00:18 songwenping_ there are extra 8 computes without gpu.
10:00:22 sean-k-mooney[m] its 13 hosts each with 8 gpus and the vm is asking for 8
10:00:53 sean-k-mooney[m] so it wont explode like that there is only 1 allocation pooible per host
10:01:01 sean-k-mooney[m] *possible
10:01:01 gibi nope
10:01:08 gibi if you have 8 groups and 8 gpus
10:01:19 gibi then each group can be satisfied by each gpu
10:01:25 sean-k-mooney[m] you think it would b n squared
10:01:41 sean-k-mooney[m] oh hum maybe
10:01:48 gibi 8!
10:02:48 gibi 40320
10:03:00 sean-k-mooney[m] ya… that would be bad
10:03:33 sean-k-mooney[m] i hope thats not whats happening or that will cause issue for pci devices in placemnt too
10:03:43 gibi this is coming from the fact that placement treats groups and RPs individually. even if we have very similar RPs and very similar groups
10:04:02 gibi I can make a test case for 8 PCI devs :)
10:04:09 gibi in nova func test
10:04:12 gibi so we can prove it
10:04:32 sean-k-mooney[m] perhaps start with 4
10:05:08 sean-k-mooney[m] i mean if its a gabbit test then sure 8
10:05:13 gibi ack. first 1 will go and try to stabilize the fair lock unit tests then I will look at the 4-8 PCI issue
10:05:45 sean-k-mooney[m] we really need to not do all possible permuations on the placement side
10:06:11 gibi if this is real factorial then we need to introduce an per host limit for a_c query
10:06:28 sean-k-mooney[m] ya i was debating if that was the only option
10:06:33 gibi if the groups are different by trait then we need all permutations
10:06:51 gibi otherwise we might not found the right one
10:07:16 sean-k-mooney[m] well an allocation candiate by defieniton meets the requirements
10:07:16 gibi the trick here is that we have identical groups and identicaly PRs
10:07:23 sean-k-mooney[m] so it depend on whyer this is happening
10:07:48 gibi to find a_c placement needs to iterate permutations I think in the general case
10:07:51 sean-k-mooney[m] anyway something to look into i guess
10:07:55 gibi yepp
10:08:36 sean-k-mooney[m] im going to grab a coffee and take my blood pressre meds and then check on freya be back in about 10 mins
10:09:02 gibi ack
10:09:22 gibi for me coffee is the blood pressure med
10:32:32 sean-k-mooney gibi: should fastener reintoudce the workaround they had
10:35:20 gibi that is a way too but I have less authority over that project
10:35:51 sean-k-mooney isnit it an oslo deliverable too
10:36:09 sean-k-mooney or has it move out of openstack
10:36:56 sean-k-mooney oh its not an openstack project
10:37:02 sean-k-mooney i tough it used to be at one point
10:38:21 sean-k-mooney i guess we can fix it in oslo but we should leth the know that under eventlet the rentrant guarentee is broken
10:38:25 sean-k-mooney https://github.com/harlowja/fasteners#-overview
10:38:35 sean-k-mooney then note that it should be reentrant
10:44:17 sean-k-mooney gibi: https://github.com/harlowja/fasteners/issues/86
10:44:17 sean-k-mooney that also intereisng
10:44:17 sean-k-mooney https://github.com/harlowja/fasteners/pull/87/files
10:44:17 sean-k-mooney elif not self.has_pending_writers:
10:44:17 sean-k-mooney elif (self._writer == me) or not self.has_pending_writers:
10:49:45 gibi I'm not sure I follow how this connects to your current problem.
10:49:52 gibi our
10:50:16 sean-k-mooney its relying on threading.current_thread
10:50:35 sean-k-mooney an now it allows you to reaquire the lock if tha tis the same
10:50:47 sean-k-mooney but with spwan_n
10:51:00 sean-k-mooney that means two greenthread coudl get the same lock if they run on the same os thread
10:51:17 gibi yes, ReaderWriterLock rely on current_thread for reentrancy

Earlier   Later