| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-11-16 | |||
| 16:41:31 | sean-k-mooney | in general you would not want to hardcode perfromace if your trying to manage power | |
| 16:41:58 | sean-k-mooney | so if the high power one was configurable that woudl be enough | |
| 16:42:15 | sean-k-mooney | if you want to keep it simple howwever why not add two config options | |
| 16:42:36 | sean-k-mooney | cpu_high_power_govoner and cpu_low_power_govoner | |
| 16:42:41 | sean-k-mooney | just simple stings | |
| 16:42:58 | bauzas | yup, that's what I write now | |
| 16:43:05 | sean-k-mooney | works for me | |
| 16:43:10 | bauzas | we could bikeshed on the namings | |
| 16:43:11 | sean-k-mooney | gibi: ^ | |
| 16:43:52 | opendevreview | Sylvain Bauza proposed openstack/nova-specs master: Proposes cpu power managment in libvirt https://review.opendev.org/c/openstack/nova-specs/+/861591 | |
| 16:43:59 | bauzas | there you go | |
| 17:10:27 | opendevreview | Sylvain Bauza proposed openstack/nova master: Deprecate mdev creation and hardfail on reboot when missing. https://review.opendev.org/c/openstack/nova/+/864418 | |
| 17:10:28 | opendevreview | Sylvain Bauza proposed openstack/nova master: Reproducer for bug 1951656 https://review.opendev.org/c/openstack/nova/+/850673 | |
| 17:10:28 | opendevreview | Sylvain Bauza proposed openstack/nova master: Handle mdev devices in libvirt 7.7+ https://review.opendev.org/c/openstack/nova/+/838976 | |
| 17:25:46 | opendevreview | John Garbutt proposed openstack/nova master: Ironic nodes with instance reserved in placement https://review.opendev.org/c/openstack/nova/+/864773 | |
| 17:28:23 | johnthetubaguy | gibi: I think I may have found an "interesting" way to close that ironic scheduler race with automatic clean, but it feels a bit wrong, its a follow on from that patch you reviewed yesterday: https://review.opendev.org/c/openstack/nova/+/864773 | |
| 17:31:34 | gibi | johnthetubaguy: that is a clever one | |
| 17:32:21 | johnthetubaguy | it might be too clever, but it seems to work... funky right? | |
| 17:33:10 | gibi | I think we intentionally designed placement in a way that it allows reserving already allocated inventories | |
| 17:33:27 | gibi | but not allow allocating reserved ones | |
| 17:33:47 | gibi | so I think what you do is OK from placement perspective | |
| 17:36:22 | gibi | how likely that a big ironic deployment has no cleaning configured and the update_provider_tree periodic is set to a big value due to preformance reasons? | |
| 17:36:35 | johnthetubaguy | OK, neat. I remember some related discussions around increasing host reserved memory, and it sounded OK ish. | |
| 17:37:14 | johnthetubaguy | ... unsure, that is a very good question. Most large deployments I work on do automatic cleaning, but I know that far from representative | |
| 17:37:38 | johnthetubaguy | (I have to run I am afraid, I am told my dinner is going cold) | |
| 17:37:59 | gibi | johnthetubaguy: no worries. I will leave the feedback in the review | |
| 17:40:42 | gibi | have a nice dinner | |
| 17:45:53 | sean-k-mooney | o/ just saw the comment on ironic reivew | |
| 17:46:08 | sean-k-mooney | is the plan to proceed with https://review.opendev.org/c/openstack/nova/+/842478 anyway | |
| 17:46:17 | sean-k-mooney | if so i can review again tomorrow | |
| 17:47:44 | sean-k-mooney | i am mostly out of brain power today but i dont mind looking at it in the morning | |
| 17:48:41 | sean-k-mooney | https://review.opendev.org/c/openstack/nova/+/864773 is the other one? to review with https://review.opendev.org/c/openstack/nova/+/842478 | |
| 17:51:52 | sean-k-mooney | and yes if you are over allcoated you cannot make new allcoations | |
| 17:52:01 | sean-k-mooney | but i think you are right that you can modify the reserve value | |
| 17:53:04 | sean-k-mooney | like you can go form reservice 0 cpus to 10 and if that would over allocate that is fine it will be resolved when a vm is delete/moved | |
| 17:54:01 | sean-k-mooney | i have not looked at johnthetubaguy clever solution but form context i assume it invovles setting the reserved value to 1 for the custom resoucs inventory that represnt the baremetal hosts | |
| 18:03:08 | opendevreview | Rafael Weingartner proposed openstack/nova master: Nova to honor "cross_az_attach" during server(VM) migrations https://review.opendev.org/c/openstack/nova/+/864760 | |
| 22:02:18 | opendevreview | Rafael Weingartner proposed openstack/nova master: Nova to honor "cross_az_attach" during server(VM) migrations https://review.opendev.org/c/openstack/nova/+/864760 | |
| 23:41:22 | opendevreview | Younghwan Yoo proposed openstack/nova master: Update return value to be a valid UUID https://review.opendev.org/c/openstack/nova/+/864674 | |
| #openstack-nova - 2022-11-17 | |||
| 01:57:53 | melwitt | bauzas: this patch looks relevant to your interests https://review.opendev.org/c/openstack/nova/+/864674 | |
| 05:59:57 | opendevreview | Jorhson Deng proposed openstack/nova master: Optimize the small pagesize in numa_fit_instance_to_host https://review.opendev.org/c/openstack/nova/+/864812 | |
| 08:04:56 | opendevreview | Jorhson Deng proposed openstack/nova master: Optimize the small pagesize in numa_fit_instance_to_host https://review.opendev.org/c/openstack/nova/+/864812 | |
| 08:42:46 | opendevreview | Jorhson Deng proposed openstack/nova master: Optimize the small pagesize in numa_fit_instance_to_host https://review.opendev.org/c/openstack/nova/+/864812 | |
| 09:44:05 | johnthetubaguy | When you run functional tests on your dev box, and you get loads of errors due to no valid host, it probably means I am doing something stupid, has anyone else hit that at all please? | |
| 10:09:15 | bauzas | johnthetubaguy: which testcases ? | |
| 10:11:16 | johnthetubaguy | good question, mostly the ones that use placement | |
| 10:11:41 | johnthetubaguy | there are loads of them in the functional suite, one example is: nova.tests.functional.wsgi.test_services.TestServicesAPI.test_resize_revert_after_deleted_source_compute | |
| 10:11:54 | johnthetubaguy | pretty sure its my environment being broken, as so many fail | |
| 10:19:47 | johnthetubaguy | ah, so I think it was by default python3 being 3.8, oops! | |
| 10:20:08 | johnthetubaguy | curiously, unit tests are just fine, but functional tests, no so much | |
| 10:22:12 | bauzas | johnthetubaguy: sorry was on a meeting | |
| 10:22:28 | johnthetubaguy | no worries, the fix was simple enough in the end | |
| 10:22:34 | bauzas | johnthetubaguy: yesterday I ran some reshape function tests locally and I had no problem | |
| 10:22:48 | johnthetubaguy | yeah, its my bad default python version, thats all | |
| 10:22:56 | bauzas | johnthetubaguy: oh, so due to py38 ? strange if so | |
| 10:23:02 | johnthetubaguy | yeah | |
| 10:23:11 | bauzas | I guess you recreated your venv too ? | |
| 10:23:31 | johnthetubaguy | I suspect something had failed to start and I didn't notice that, all I noticed was lots of no valid host exceptions | |
| 10:23:43 | johnthetubaguy | so tox spotted it and did the re-create | |
| 10:23:56 | bauzas | and it worked then ? | |
| 10:24:13 | johnthetubaguy | in tox.ini the base python is python3 rather than python3.9 I guess, which caused my "fun" | |
| 10:24:18 | johnthetubaguy | yeah, its working fine now | |
| 10:24:21 | bauzas | I suspect some os-resource-classes or os-traits library update wasn't updated | |
| 10:24:35 | bauzas | hence placement failing and then the novalidhosts eventually | |
| 10:24:45 | bauzas | doing the tox -r updated the deps | |
| 10:24:47 | johnthetubaguy | ah... got you, very possible | |
| 10:25:10 | johnthetubaguy | actually, I remember seeing some os-traits error actually | |
| 10:25:20 | bauzas | yeah, we hardfail on the number of traits | |
| 10:25:40 | johnthetubaguy | it was a missing attribute error I think, which is basically the same thing | |
| 10:25:48 | bauzas | cool | |
| 10:25:58 | johnthetubaguy | well mystery solved, thank you! | |
| 10:26:07 | bauzas | np, glad you fixed it by yourself :D | |
| 10:26:31 | johnthetubaguy | (makes bashing with a hammer noises) | |
| 10:27:06 | bauzas | :) | |
| 10:44:00 | opendevreview | John Garbutt proposed openstack/nova master: Functional test test_boot_reschedule_with_proper_pci_device_count https://review.opendev.org/c/openstack/nova/+/760354 | |
| 10:44:01 | opendevreview | John Garbutt proposed openstack/nova master: Fix PCI passthrough race on reschedule (refresh) https://review.opendev.org/c/openstack/nova/+/710848 | |
| 10:44:01 | opendevreview | John Garbutt proposed openstack/nova master: Fix PCI passthrough race on reschedule (claims) https://review.opendev.org/c/openstack/nova/+/710847 | |
| 11:02:59 | johnthetubaguy | gibi: I think you reviewed those in the past, I am not sure it answers all your questions, but I put the functional test first, in an attempt to work out which patches are needed. Its a nasty bug that on re-schedule you try to get the wrong PCI device, but fail with in-use errors. | |
| 11:03:54 | sean-k-mooney | johnthetubaguy: so i was thinking about your ironic patch for reserving on schedule over night | |
| 11:04:04 | sean-k-mooney | i think in general its a good idea | |
| 11:04:32 | sean-k-mooney | there was some concern about if cleaning was used or not in large cloud right extendign the time they would be unavaiable | |
| 11:05:04 | johnthetubaguy | yeah, gibi was mentioning that, and well, I don't disagree | |
| 11:05:11 | sean-k-mooney | if we wanted to cater for that we coudl make this configurable but i think the its proably ok ot reserve by default or uncondtionally | |
| 11:05:39 | johnthetubaguy | yeah, it feels like a future workaround config, if its a problem for people | |
| 11:05:56 | johnthetubaguy | interestingly, I think it fixes an extra case I should add to the commit... | |
| 11:06:25 | sean-k-mooney | oh what one | |
| 11:07:12 | johnthetubaguy | when you mark an in-use node as in maintenance mode as its broken, user gets to delete their instance when they are ready, and that goes into clean failed (depending on your ironic config), we don't hit the race with our placement updates any more either | |
| 11:07:27 | johnthetubaguy | we had the same window with that, once the allocation is removed when the instance is deleted | |
| 11:08:05 | sean-k-mooney | oh ok so clean failed happens because it in mantainance | |
| 11:08:14 | sean-k-mooney | and cant actully start cleaning? | |
| 11:08:55 | johnthetubaguy | more that, we start sending new instances to the node that is in maintainance, shortly after the user deletes their nova server | |
| 11:09:26 | johnthetubaguy | its basically the same race condition, but with a slightly different reason | |
| 11:09:45 | sean-k-mooney | nice i alwasys like it when one fix fixes multiple bugs | |
| 11:10:13 | johnthetubaguy | totally | |
| 11:11:10 | sean-k-mooney | so between https://review.opendev.org/c/openstack/nova/+/842478 (the retry) and https://review.opendev.org/c/openstack/nova/+/864773 (reserving) we have two fixes. they retry is certinaly backportable | |
| 11:11:33 | sean-k-mooney | reserving honelsy proably is too but to backport that i think we would need the workaround option | |
| 11:12:05 | johnthetubaguy | yeah, I think we need both, although the new one makes the older one less important | |
| 11:12:55 | johnthetubaguy | i.e. the older one only matters where available nodes go no longer available, and we don't spot it right away, since we remove the issue with automatic cleaning also causing that problem | |
| 11:14:12 | sean-k-mooney | yes so i was about to approve the old patch and then soft -1 the second one askign for the workaround option if that works for you. | |
| 11:14:33 | sean-k-mooney | the only thing i was wonderign about for the first patch is shoudl it have a release note | |