| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2021-11-11 | |||
| 11:38:16 | noonedeadpunk | well I wasn't able to get it working tbh | |
| 11:38:29 | sean-k-mooney | in this case each VF will expose a singel mdev | |
| 11:38:45 | noonedeadpunk | so I just don't have mdev_bus until I run /usr/lib/nvidia/sriov-manage -e ALL | |
| 11:38:45 | sean-k-mooney | what nvidia does not support is using the VF via pci pasthough directly | |
| 11:39:03 | sean-k-mooney | yes | |
| 11:39:12 | sean-k-mooney | that i quirk of the driver | |
| 11:39:42 | noonedeadpunk | aha, ok. I'm jsut a bit confused by amount of pci devices | |
| 11:39:43 | sean-k-mooney | noonedeadpunk: bauzas is on PTO today as its a holidy in france but they have been doing some early investigation with it have booted vms | |
| 11:39:57 | sean-k-mooney | let me see if i can get there write up | |
| 11:41:30 | noonedeadpunk | Eventually I have A10s atm, and if to follow https://docs.nvidia.com/grid/latest/grid-vgpu-user-guide/index.html#vgpu-types-nvidia-a10 I can get 6 vgpus of A10-4C for example. But I can hardly understand how placement would pick 6 out of these all pci devices considering every has this mdev type | |
| 11:41:53 | sean-k-mooney | https://bugzilla.redhat.com/show_bug.cgi?id=1975872#c2 is tracking our docs update for how to configure it | |
| 11:42:04 | sean-k-mooney | oh that is private hum | |
| 11:43:25 | noonedeadpunk | like https://paste.opendev.org/show/810935/ which does not make any sense to me... | |
| 11:43:57 | sean-k-mooney | why does that look odd to you | |
| 11:43:58 | noonedeadpunk | and ofc nvidia support tries their best to be not helpful... | |
| 11:44:30 | noonedeadpunk | um, because I can create only 6 vgpus of such type? | |
| 11:44:50 | noonedeadpunk | but the way more devices are repoted to placement | |
| 11:45:27 | sean-k-mooney | right ok so the nvidia smi tool basically seams to ignore the count you ask for i think that was one fo the quirks bauzas mentioned | |
| 11:45:47 | sean-k-mooney | noonedeadpunk: i think yhou will need to whitelist only a subset of the devices | |
| 11:45:52 | noonedeadpunk | you actually can't even provide to it anything | |
| 11:45:58 | noonedeadpunk | it has jsut enable or disable... | |
| 11:48:17 | noonedeadpunk | ok, thanks anyway for pointers - I will return back tomorrow I guess to catch bauzas :) | |
| 12:47:08 | lyarwood | gibi / sean-k-mooney ; finally found the change in tempest I've been talking about this week btw https://github.com/openstack/tempest/commit/e3405ba808f97eae57f3a60991000afaa34cbe89 | |
| 12:47:30 | lyarwood | wait_for_sshable=True will read console output until it see's a login: prompt | |
| 12:47:35 | lyarwood | kinda awful | |
| 12:47:40 | gibi | lyarwood: ohh | |
| 12:47:50 | gibi | lyarwood: how much slower it will make our test run? | |
| 12:48:02 | gibi | lyarwood: but anyhow, it worth to try at least | |
| 12:48:12 | lyarwood | I'm not sure tbh but I can't see how that's a valid thing to enable in Tempest | |
| 12:48:30 | lyarwood | like what says that a test image has to contain a guestOS that will actually print that? | |
| 12:49:01 | lyarwood | or configurable per tempest run and not per test | |
| 12:49:35 | sean-k-mooney | lyarwood: why | |
| 12:49:42 | sean-k-mooney | but ok i see why that is slow | |
| 12:50:02 | sean-k-mooney | it really shoudl jsut try to ssh in a retry loop | |
| 12:50:02 | lyarwood | well there's nothing stopping us from using a Windows image during the test run right? | |
| 12:50:13 | sean-k-mooney | like the pingable one did | |
| 12:50:14 | lyarwood | and that wouldn't print a login prompy | |
| 12:50:22 | lyarwood | prompt* | |
| 12:50:47 | sean-k-mooney | yes but for windows we woudl also use winrm not ssh | |
| 12:50:52 | lyarwood | and my point about making this configurable per run as opposed to per test is because you would have to go in and add this kwarg to every single create_test_server call | |
| 12:51:07 | gibi | /o\ | |
| 12:51:12 | lyarwood | sean-k-mooney: right but the change I've pointed to is reading console output | |
| 12:51:23 | sean-k-mooney | ya which i think is wrong | |
| 12:51:48 | lyarwood | Wonderful, just making sure | |
| 12:52:30 | sean-k-mooney | so the old validation code support waithign for the server to be reacable via ping or ssh via config | |
| 12:52:40 | sean-k-mooney | i was expecting it to jsut use that | |
| 12:52:54 | sean-k-mooney | which did not have any console interaction if i understand correctly | |
| 12:53:28 | lyarwood | There's nothing I can see in the create code that did this previously | |
| 12:53:43 | lyarwood | AFAICT this was added as a step prior to the tests attempting to SSH into the instance | |
| 12:53:55 | lyarwood | to ensure the instance had booted up and was ssh'able itself | |
| 12:57:07 | sean-k-mooney | ack so i tought you can use https://github.com/openstack/tempest/blob/master/tempest/config.py#L897-L970 to configure vlaidations for all test | |
| 12:57:22 | sean-k-mooney | can we use that to adress this in the job config | |
| 12:58:13 | lyarwood | that doesn't actually do anything for most tests | |
| 12:58:27 | lyarwood | brb need to jump on a call | |
| 12:59:16 | sean-k-mooney | ... i tought this was ment to run on any test that created a server as part of the server create automatically | |
| 13:06:30 | gibi | sean-k-mooney: https://github.com/openstack/tempest/blob/master/tempest/scenario/manager.py#L229-L236 I think this describes the situation | |
| 13:06:55 | gibi | sean-k-mooney: so there was an intention to allow running validation for each server but it was never introduced globally | |
| 13:11:38 | sean-k-mooney | right i remember that being the intent | |
| 13:11:57 | sean-k-mooney | and in the past i think you could even confiure that validation metion to be either ping or ssh | |
| 13:12:09 | sean-k-mooney | there was a spec for this somewhere | |
| 13:15:14 | sean-k-mooney | this https://specs.openstack.org/openstack/qa-specs/specs/tempest/implemented/ssh-auth-strategy.html | |
| 13:15:27 | sean-k-mooney | """ it extends the valid value for wait_until with new types of wait abilities: PINGABLE and SSHABLE. """ | |
| 13:17:27 | sean-k-mooney | ... https://github.com/openstack/tempest/blob/ed89c77222917235290c8cc51974835528ed4cfa/tempest/common/compute.py#L101 | |
| 13:17:55 | sean-k-mooney | so ya i twas not actully implemeted | |
| 13:18:05 | gibi | yeah | |
| 13:18:19 | sean-k-mooney | this is also not the first time i have wanted to use this and discoverd this | |
| 13:18:26 | gibi | :) | |
| 13:19:44 | sean-k-mooney | maybe we shoudl jsut implement it | |
| 13:19:50 | sean-k-mooney | at least the pingable version | |
| 13:20:16 | sean-k-mooney | ssh would be nice but if it has an ip the os should be live enough for hotplug | |
| 13:20:25 | gibi | I think the comment also states that pingability means some level of network setup is in place and from the create_server perspective this cannot be ensured | |
| 13:20:29 | lyarwood | Yup I can hack on this | |
| 13:21:01 | sean-k-mooney | gibi: well we woudl wait for active and then waith for pingable | |
| 13:21:10 | sean-k-mooney | only if the server has a network | |
| 13:21:24 | gibi | do we need floating ip for pingability? | |
| 13:21:30 | lyarwood | has a network, fip etc yeah | |
| 13:21:31 | gibi | and securty group setuo? | |
| 13:21:45 | lyarwood | yeah we'd need to pass in and setup the validation resources | |
| 13:21:58 | lyarwood | it's entirely possible in the base server creation method | |
| 13:22:33 | gibi | OK then we are on the same page about what is required | |
| 13:22:51 | lyarwood | I'll revert the console stuff and try to hack on this in the background of calls this afternoon | |
| 13:23:08 | gibi | thank you lyarwood | |
| 13:23:30 | sean-k-mooney | by the way if we just put a sleep(300) in the test will the issue go away | |
| 13:23:55 | gibi | sean-k-mooney: that can be tried too | |
| 13:24:09 | sean-k-mooney | e.g. before you do all that work which is good, are we confident it will help | |
| 13:24:18 | lyarwood | yup that's a fair test | |
| 13:24:21 | lyarwood | maybe not 300 | |
| 13:24:23 | sean-k-mooney | i think it might if its an issue with the guest not being ready | |
| 13:24:29 | sean-k-mooney | well ya mayb like 30 | |
| 13:24:30 | lyarwood | as other things will likely timeout | |
| 13:24:47 | lyarwood | okay if someone can test that it would be great | |
| 13:25:09 | gibi | I will push a tempest patch and a nova depends-on for that | |
| 13:25:17 | gibi | * for the sleep casse | |
| 13:25:18 | gibi | case | |
| 13:25:53 | sean-k-mooney | how is it already half 1 | |
| 13:26:19 | sean-k-mooney | not that the conversation is not engagin but i keep getting distracted today | |
| 13:26:41 | sean-k-mooney | i ment ot start with the off path acclerator spec this morning | |
| 13:27:58 | gibi | I feel your pain sean-k-mooney I had a good day on tuesday but wednesday was a loss and today doesn't look good either :) | |
| 13:40:55 | gibi | lyarwood sean-k-mooney: so what I see is that test_live_block_migration_with_attached_volume causing the most kernel panic (if not all) and the panic happens when tempest runs the resource cleanup after the whole test class. So I will add the extra sleep at the top of the volume detach code to see if that helps | |
| 13:41:20 | gibi | does it sounds good to you? | |