| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-09-28 | |||
| 13:17:40 | opendevreview | Justas Poderys proposed openstack/nova-specs master: Add support for Napatech LinkVirt SmartNICs https://review.opendev.org/c/openstack/nova-specs/+/859290 | |
| 13:18:57 | justas_napa | Hi bauzas, gibi. Re: Napatech smartnic support, we have now pushed the code changes to opendev | |
| 13:19:22 | justas_napa | all changes are visible in opendev via https://review.opendev.org/q/topic:Napatech_SmartNIC_support | |
| 13:45:05 | noonedeadpunk | gibi: I guess the question is - why new allocation is created at the first place. As it's already present, isn't it? | |
| 13:46:20 | noonedeadpunk | so why another allocation is being created against source host | |
| 13:47:36 | gibi | noonedeadpunk: it is a technicality. When the instance is created the resource allocation of the instance is held by the consumer=instance.uuid in Placemnet. When the instance is being migrated, Nova first moves the allocation held by the consumer=instance.uuid to consumer=migration.uuid in Placement, then nova will call the scheduler to select a target host for the migration and allocate the | |
| 13:47:42 | gibi | resource on the target host to consumer=instance.uuid in Placement. | |
| 13:48:21 | gibi | noonedeadpunk: the move of the allocation from instance.uuid to migration.uuid is the one that considered as a new allocation in Placement and rejected as the host is already overallocated | |
| 13:48:50 | noonedeadpunk | well, it's quite slight line between bug and "as designed" | |
| 13:49:40 | noonedeadpunk | because for me - allocation change that does not involve change of resources shouldn't be considered as new one... | |
| 13:49:44 | gibi | I intentionally not say it is a bug or not a bug, what I say is that having the host overallocated (over the overallocation limit) is the cause of the problem. A host should never be in that state | |
| 13:50:54 | noonedeadpunk | but I do agree that workaround is easy enough to not spend time on making placement logic mroe complex for this | |
| 13:51:12 | gibi | in all these cases you have to first figure out why and how you ended up in an overallocated state | |
| 13:51:36 | gibi | and as you said recovering from this state is not as hard | |
| 13:52:37 | noonedeadpunk | I can say super simply example - you decided that overcommit ratio is too high and it's good to lower it down. And you have prometheus that can tell ou on average for region overcommit is lower then you want to define | |
| 13:53:24 | noonedeadpunk | So you lower it down, but some computes are still having higher overcommit, as you have not checked that against placement per host | |
| 13:53:46 | noonedeadpunk | And you can't lower it down easily by reverting setting and disabling compute for scheduling on top | |
| 13:55:17 | noonedeadpunk | * And you can't lower it down easily by migrating instances out so you have to revert setting | |
| 13:59:58 | gibi | basically that process makes the mistake of lowering the allocation ratio on the host before checking it if it causes overallocation | |
| 14:01:41 | gibi | so the proper process would be (in my eyes): 1) disable the host temporarily to avoid new scheduling while you reconfigure it 2) migrate VMs out from the host to achive the target resource usage 3) reconfigure the host with the new allocation ratio 4) enable the host | |
| 14:03:00 | noonedeadpunk | well... it's hard to disagree here :) | |
| 14:03:34 | noonedeadpunk | (it's more tricky to do though with microversion 2.88) | |
| 14:07:51 | noonedeadpunk | or maybe not and it's jsut me who got used to deal with nova api... | |
| 14:08:36 | noonedeadpunk | as eventually representation of inventory in placement is way more accurate | |
| 14:10:12 | gibi | yeah you should use placement to calculate how many vms you need to move | |
| 14:29:26 | auniyal | Hello #openstack-nova | |
| 14:29:42 | noonedeadpunk | gibi: the problem is that with sdk it's quite tricky to use placement as you need to really call api rather then get resources with all attributes in it (like it was with hypervisors) | |
| 14:30:07 | auniyal | how this works - https://opendev.org/openstack/nova/src/commit/aad31e6ba489f720f5bdc765c132fd0f059a0329/nova/context.py#L396 | |
| 14:55:54 | noonedeadpunk | gibi: just to express my pain as operator. https://paste.openstack.org/show/bpA2Iq28mHEfo8r4sH2W/ consumes about 3 seconds to execute. And as you see - overrides microversion <2.88. | |
| 14:56:13 | noonedeadpunk | If I am to use palcement to get the same data - I will need code like that https://paste.openstack.org/show/bOsJwmnMMpftgZ01St0o/ and it takes... 42 seconds to execute against exact same cluster | |
| 14:57:28 | noonedeadpunk | And not saying about load on APIs as for each resource provider I will need to make 2 calls to placement | |
| 14:59:41 | noonedeadpunk | so for users deprecation of providing resource statistics from hypervisors is quite serious regression in performance. And if I want to monitor that regulary - it's really waste of time. | |
| 15:00:02 | gibi | noonedeadpunk: interesting. so there is some heavy inefficiencies somewhere as it just 2x the number of calls but it takes more than 10 times the time to execute | |
| 15:00:49 | noonedeadpunk | fwiw - I was testing remote cloud (not from inside of it) | |
| 15:03:28 | noonedeadpunk | just to show that I'm not exaggerating https://paste.openstack.org/show/b1irA4GLwJEdUORuLarb/ | |
| 15:04:40 | noonedeadpunk | But I'm not sure about 2x number of calls. As it feels that from hypervisor data was fetched with single api call... | |
| 15:04:47 | noonedeadpunk | Not sure though | |
| 15:05:04 | noonedeadpunk | (or at least I can't explain it otherwise) | |
| 15:06:31 | gibi | hm, you are right the hypervisor one returned all compute from the cluster | |
| 15:06:34 | gibi | /os-hypervisors/detail | |
| 15:06:46 | gibi | so then I think it would make sense to add a similar api for placement | |
| 15:06:55 | gibi | single call that returns all the usages per provider | |
| 15:07:28 | gibi | this explains the preformance differences | |
| 15:08:00 | noonedeadpunk | yeah and it will depend on amount of providers actually right now quite dramatically | |
| 15:08:57 | gibi | so I would support adding such api to placement. I cannot commit to implement it, but sure I can review the implementation | |
| 15:09:46 | noonedeadpunk | I can only put it to my backlog and hope to get to it one day... | |
| 15:10:02 | gibi | noonedeadpunk: and if you want to raise this issue then I think we have dedicated time on the coming PTG for operator feedback | |
| 15:10:22 | noonedeadpunk | that is good idea actually | |
| 15:11:43 | gibi | noonedeadpunk: https://etherpad.opendev.org/p/oct2022-ptg-operator-hour-nova | |
| 15:11:44 | noonedeadpunk | first hour does not intersect with anything, so can join:) | |
| 15:11:52 | gibi | cool | |
| 15:17:23 | jkulik | you couldâ„¢ make the calls in parallel at least ;) | |
| 15:19:07 | noonedeadpunk | isn't it still waste of resources? | |
| 15:19:24 | noonedeadpunk | and regression from operator prespective? | |
| 15:19:43 | jkulik | sure. would be helpful to get it in a single call, but at least it reduces the pain of waiting | |
| 15:20:06 | noonedeadpunk | and add some hardware to serve increased load?:) | |
| 15:20:23 | noonedeadpunk | but yeah, that's fair | |
| 15:20:35 | noonedeadpunk | *fair workaround | |
| 15:28:40 | noonedeadpunk | fwiw, even with 10 threads it's still 5 times slower then jsut ask nova | |
| 15:28:58 | noonedeadpunk | maybe I shouldn't have used joblib... | |
| 22:07:58 | rm_work | hey, i've noticed the OSC seems to sometimes hide the "fault" field on a server show command... but also sometimes it doesn't. anyone know why this is the case? | |
| 22:08:35 | rm_work | starting to dig into the code now, but hoping someone has some idea, since the clients are often a bit funky to interpret, lots of magic 😛 | |
| 22:21:00 | rm_work | yeah, didn't find anything specifically relating to "fault" in there... this is super weird, I can see the "fault" come back with `--debug` in the response from nova, it just gets filtered out or something before the results are shown | |
| 22:31:09 | rm_work | nevermind, figured it out, there's a client patch I didn't know about >_< FML | |
| 22:35:45 | clarkb | rm_work: don't leave us hanging. What causes it? | |
| 22:41:40 | rm_work | local client patch where someone did something ... ill advised | |
| 22:41:54 | rm_work | trying to figure out how to untangle it now | |
| 22:45:30 | clarkb | Oh I see what you mean by client patch now. I thought you meant you found the chagne in gerrit taht did it or something | |
| #openstack-nova - 2022-09-29 | |||
| 08:29:10 | songwenping | bauzas: hi, i have some doubt about nova zed highlight, what is the improvement for 'the possibility to rebuild a volume-backed instance'? | |
| 08:32:14 | gibi | songwenping: this is the code we merged to suppot it https://review.opendev.org/q/topic:bp/volume-backed-server-rebuild+status:merged | |
| 08:33:46 | songwenping | is this the API imporvement? | |
| 08:39:00 | songwenping | gibi: i means this feature is not only for API, but include compute modify. | |
| 08:39:33 | gibi | it includes everythin (cinder and nova impacts) that allows reimaging the root volume of a VM booted from that volume | |
| 08:40:08 | gibi | so yes, it is not just an API change | |
| 08:42:19 | songwenping | thanks, got it. and what's the feature for 'the change to only accept importing a public key but also with an extended name pattern'. | |
| 08:44:52 | gibi | that is actually two features. 1) nova stoped supporting generating ssh keys, we only support importing ssh keys. (it is in a new microversion so old microversions still allow generating ssh keys) 2) we extended the list of allowed chars set for the name of the key. Now both @ and . is allowed in the name | |
| 08:52:55 | songwenping | thanks gibi, got it. | |
| 16:01:51 | open10k8s | Hi team, Task_state=deleting VMs don't exist on hypervisors but openstack keeps those list as Active and running. Nova-compute restart solves this but it is only a temporary solution. After delete another machine, same thing happens. I cannot find any error logs from nova services. Is this known issue in a specific condition or any workaround for this? | |
| 16:03:49 | open10k8s | I am running wallaby and never experienced this issue on other clouds with same version | |
| 16:09:02 | melwitt | open10k8s: no, that's not a known issue and it shouldn't be doing that. you are seeing this happen with all delete requests? if the actual VM is gone but the instance still listed, that means that somehow the delete did not reach the DB update to remove the instance record | |
| 16:09:35 | open10k8s | yes, all deletion operation | |
| 16:10:36 | melwitt | I would expect to see some evidence of why in the nova-compute and or nova-conductor logs. are you running logs with debug=True to diagnose? if there's no message in the logs, it will be difficult to find why it's happening | |
| 16:10:48 | open10k8s | i rather tried to find any issues on oslo messaging | |
| 16:11:39 | melwitt | yeah that is a good thing to look for, also the database if there's any issue there when nova-compute attempts to delete the instance record by way of the nova-conductor | |
| 16:12:00 | melwitt | but in both of those cases usually there is an error logged | |
| 19:33:15 | opendevreview | Ghanshyam proposed openstack/os-vif stable/stein: Fix zuul config error for os-vif https://review.opendev.org/c/openstack/os-vif/+/859892 | |
| 19:38:52 | opendevreview | Ghanshyam proposed openstack/os-vif stable/rocky: Fix zuul config error for os-vif https://review.opendev.org/c/openstack/os-vif/+/859893 | |
| #openstack-nova - 2022-09-30 | |||
| 06:48:35 | opendevreview | Amit Uniyal proposed openstack/nova master: Adds check if resized to swap zero https://review.opendev.org/c/openstack/nova/+/857339 | |
| 07:01:44 | amorin | hey nova team, if you have time someday to review this: https://review.opendev.org/c/openstack/nova/+/853682 | |
| 07:15:18 | Uggla | Poke bauzas --> https://photos.app.goo.gl/MRmeetQr749GFh3s5 ;) | |
| 07:31:11 | bauzas | danm fucking numpy | |
| 07:31:22 | bauzas | I should have prepared better | |
| 07:32:02 | bauzas | and now I have to wait until Feb 2023 IIRC | |
| 07:32:05 | bauzas | graaaah | |
| 07:34:23 | bauzas | oh no, actually after Mar 29th :( | |
| 09:15:42 | Uggla | gibi, bauzas could you have a look at https://review.opendev.org/c/openstack/nova/+/854355 | |
| 09:17:21 | bauzas | ouch. | |
| 09:17:50 | bauzas | Uggla: what do you mean by "soft delete is deprecated ?" | |
| 09:18:40 | Uggla | bauzas, This is mentioned and checked in a test that we should not create soft delete objects. | |
| 09:19:27 | bauzas | I'd say this isn't recommended, true | |