| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-12-16 | |||
| 12:04:14 | sean-k-mooney | stephenfin: im hoping to finish reviewign the rest of the pci seriese today but if not it will be my goal to get it completed the first week of january when im back | |
| 12:42:49 | songwenping | sean-k-mooney,gibi: it seems not reasonable if live migration uses 2 seperate allocations, the vm cannot be migrate if the source node have not enough resouce. | |
| 12:50:46 | sean-k-mooney | songwenping: it can we wont stop the migration if the source is over commited | |
| 12:51:20 | sean-k-mooney | that will fail for evacuate but live migration shoudl work | |
| 12:51:47 | sean-k-mooney | that said you shoudl never get into that situration unless you change something in a way that was unsupproted or hit a bug | |
| 12:52:13 | sean-k-mooney | for example reduced the memory in the source node either intentiollay or due to a dim failure or change the allcoation ratio | |
| 12:52:50 | songwenping | we use the Rocky code, and placement is integrated with nova. | |
| 12:53:17 | sean-k-mooney | multiple allocation was intoduced in Queens | |
| 12:53:23 | sean-k-mooney | so it shoudl be there in rocky | |
| 12:54:16 | songwenping | but we encounter the problem, the vm cannot migrate because the source node's memory over commit | |
| 12:55:12 | sean-k-mooney | we may have a bug in rocky or master then but you shoudl not be in an over commit senario | |
| 12:55:36 | sean-k-mooney | its fine to oversubscibe provided the total amoutn adn allocation ratio align | |
| 12:55:56 | sean-k-mooney | the temporay fix woudl be to increase the allocation ratio to enabel the migration | |
| 12:56:03 | sean-k-mooney | and then restore it when done | |
| 12:56:16 | sean-k-mooney | althernitivly you coudl try cold migration. | |
| 12:58:40 | songwenping | yes, we try to increase the allocation ratio now, and analyse why the memory is over commit | |
| 13:12:26 | gibi | yepp it is a know behavior that you cannot migrate from an overallocated node. remove the overallocation by temporarily increasing allocation ratio, move the instance, restore the allocation ratio. And separately investigate how you ended up in an overallocated scenario | |
| 13:18:45 | sean-k-mooney | gibi: i tought that only affected migrtions that uses a single allcoation | |
| 13:20:25 | gibi | sean-k-mooney: I think it is the other way around. Evacuation is the only move that does not use migration allocation on the source. It only extend the instance allocation to cover both the source and the dest node. So I think evac is not effect by this | |
| 13:20:49 | gibi | all the other move operators move the instance allocation from the instance_uuid to the migration_uuid on the source node | |
| 13:21:01 | gibi | that move allocation is what triggers the situation | |
| 13:21:10 | gibi | becuase placement does not have a move semantic | |
| 13:22:38 | sean-k-mooney | hum maybe | |
| 13:23:03 | gibi | https://github.com/openstack/nova/blob/36091a7ed7ad553d5cbb5dcfde0090e1e762bc34/nova/scheduler/client/report.py#L2023-L2043 | |
| 13:23:04 | sean-k-mooney | evacuate is broken for overcommited case too however | |
| 13:23:23 | sean-k-mooney | the extention failes if the souce is over allcoated | |
| 13:23:58 | gibi | OK then I was mistaken on that part. then all allocation manipulation is reject by placement if there is at least one overallocated RP in the change | |
| 13:24:20 | sean-k-mooney | proably yes | |
| 13:24:47 | gibi | for some reason I assumed that if the allocation on the overallocated RP does not change then it is OK and evac only adds new alloc on new RPs. | |
| 13:24:51 | sean-k-mooney | i filed a downstrema tracker for this and its a know issue upstream so hopefully this will eventually get resovled | |
| 13:24:59 | sean-k-mooney | although it might require placemtn changes | |
| 13:25:46 | gibi | the resolution probably needs a new placement microversion either to change the semantic of the existing POST /allocations or to add a new API that understands move semantic | |
| 13:26:06 | gibi | but I agree to do something as it is a common issue from deployers | |
| 13:40:45 | opendevreview | Jorge San Emeterio proposed openstack/nova-specs master: Review usage of oslo-privsep library on Nova https://review.opendev.org/c/openstack/nova-specs/+/865432 | |
| 13:49:24 | opendevreview | Ruby Loo proposed openstack/nova stable/yoga: Ironic nodes with instance reserved in placement https://review.opendev.org/c/openstack/nova/+/867912 | |
| 14:03:34 | opendevreview | Ruby Loo proposed openstack/nova stable/xena: Ironic nodes with instance reserved in placement https://review.opendev.org/c/openstack/nova/+/867913 | |
| 14:07:59 | opendevreview | Ruby Loo proposed openstack/nova stable/wallaby: Ironic nodes with instance reserved in placement https://review.opendev.org/c/openstack/nova/+/867914 | |
| 14:08:20 | gibi | bauzas, sean-k-mooney: fyi I filed two bugs in the last two days about gate instabilties as I'm hitting them https://bugs.launchpad.net/glance/+bug/1999800 https://bugs.launchpad.net/tempest/+bug/1999893 | |
| 14:09:03 | bauzas | shitty shit | |
| 14:09:37 | bauzas | gibi: thanks | |
| 14:09:59 | gibi | these are infrequent ones but I see both more than once so I reported them | |
| 14:52:05 | gibi | I'm a magnet of bugs these days | |
| 14:52:11 | gibi | the latest, this is from my local env | |
| 14:52:12 | gibi | Dec 16 15:51:21 bedrock kernel: traps: flake8[1268565] general protection fault ip:55f5a5ed6e83 sp:7ffdf4a39d50 error:0 in python3.10[55f5a5dba000+2a3000] | |
| 14:52:35 | gibi | I cannot even run tox -e pep8 as flake8 fails all the time | |
| 14:54:33 | ykarel | Hi can someone look into https://bugs.launchpad.net/nova/+bug/1949606 | |
| 14:55:02 | ykarel | libvirt-8.0.0 now provides option to set tb-cache | |
| 14:56:03 | ykarel | without it it's difficult to run multiple guests vm together in CI on jammy hosts | |
| 15:07:27 | gibi | ykarel: can we default tb-cache size globally via some libvir configuration? | |
| 15:07:30 | gibi | kashyap: ^^ | |
| 15:08:17 | ykarel | gibi, no idea, but if that's possible then would be helpful as can be set outside of nova too | |
| 15:08:52 | gibi | ykarel: yep, it would be convinient otherwise we need to create a nova feature just for our CI usage | |
| 15:09:13 | ykarel | yes | |
| 15:09:48 | gibi | i.e. a new nova compute host level config variable in [libvirt] section set to some small value applied blindly to all emulated domains by the nova-compute service | |
| 15:10:30 | ykarel | it used to be 32MiB before it was raised to 1GiB | |
| 15:10:46 | ykarel | so that should be good for CI atleast | |
| 15:13:51 | kashyap | gibi: Hmmm, good question | |
| 15:14:12 | kashyap | gibi: It rings a faint bell as I looked at it in the past, but I forget | |
| 15:14:36 | kashyap | I'm in a hurry as I need to take a train shortly, but I'll take a quick look | |
| 15:14:49 | gibi | kashyap: no worries, it is not super urgent :) | |
| 15:15:05 | kashyap | gibi: Good news: yes! libvirt does allow it | |
| 15:15:22 | kashyap | LOL, I tested it even upstream libvirt myself and totally forgot: | |
| 15:15:55 | kashyap | gibi: ykarel: https://listman.redhat.com/archives/libvir-list/2021-November/224873.html | |
| 15:16:43 | ykarel | kashyap, yeap i tested that and it works, now we are looking if we can set it globally by some libvirt conf | |
| 15:16:53 | gibi | kashyap: with my limited understanding it only show that it is allowed via the domain xml, can we also set it via some hypervisor level global config? | |
| 15:16:55 | ykarel | so we don't have to change nova code just to support CI usecase | |
| 15:19:38 | kashyap | gibi: ykarel: I don't think global config is possible - near as I know | |
| 15:19:47 | gibi | kashyap: thanks | |
| 15:19:48 | kashyap | ykarel: But just shoot an email to libvirt-users@redhat.com list and ask there. | |
| 15:19:53 | kashyap | People are friendly :) | |
| 15:20:25 | ykarel | kashyap, Ok Thanks | |
| 15:20:35 | ykarel | will send a mail | |
| 15:23:33 | kashyap | ykarel: A quick tip: Ask them to keep you explicitly in Cc you on responses, as you're not subscribed to that list (I guess) | |
| 15:24:00 | opendevreview | Balazs Gibizer proposed openstack/nova master: Split ignored_tags in stats.py https://review.opendev.org/c/openstack/nova/+/867978 | |
| 15:24:14 | gibi | sean-k-mooney: I did the split as we discussed ^^ | |
| 15:24:37 | ykarel | Thanks kashyap, yes right /me not subscribed | |
| 16:17:46 | ykarel | kashyap, gibi sent https://listman.redhat.com/archives/libvirt-users/2022-December/013844.html | |
| 16:26:29 | rloo | hi sean-k-mooney, these should be ready to approve (zuul is happy anyway!): https://review.opendev.org/c/openstack/nova/+/867912, https://review.opendev.org/c/openstack/nova/+/867913 & https://review.opendev.org/c/openstack/nova/+/867914 (thanks!) | |
| 17:18:17 | sean-k-mooney | rloo: ack | |
| 17:29:16 | opendevreview | Edward Hope-Morley proposed openstack/nova stable/yoga: ignore deleted server groups in validation https://review.opendev.org/c/openstack/nova/+/867989 | |
| 17:31:50 | sean-k-mooney | rloo: the older backports are not quite right | |
| 17:32:14 | rloo | sean-k-mooney: gahhhh. did you comment? I'll take a look. | |
| 17:32:15 | sean-k-mooney | the content is fine but it looks like you cherry picked form master in all cases instead of form the previous cherry pick | |
| 17:32:29 | sean-k-mooney | yep its pretty minor | |
| 17:32:38 | rloo | yes, i cherry picked from master. do you want to do it from previous cherry pick? | |
| 17:32:39 | sean-k-mooney | just the commit is wrong | |
| 17:33:05 | sean-k-mooney | rloo: yep you should cherry pick form the previous cherry prick | |
| 17:33:35 | sean-k-mooney | i think ironic does this slightly differntly due to how ye do bugfix branches | |
| 17:33:47 | rloo | geez. i thought if i used the UI to do the cherry pick, it'd do the right thing. the reason i didn't do from previous, was cuz things looked messier, heh. | |
| 17:33:57 | sean-k-mooney | for nova the backport go form newest to oldest branch and you cherry pick form the previosu branch | |
| 17:34:25 | rloo | i haven't been doing upstream stuff, so i don't even recall how ironic does it... i did try to find doc about it but gave up. | |
| 17:34:46 | sean-k-mooney | rloo: i acttully care about the cherry-pick lines less then other but i knwo melwitt and elodilles do like them to be done a specific way | |
| 17:35:41 | sean-k-mooney | for me i just do a git reset --hard origin/stable/<whatever> then git review -X <previous version> | |
| 17:35:51 | rloo | no worries. should i create new PRs, the 'right' way? | |
| 17:36:39 | sean-k-mooney | well they dont have to be new reviews just need to fix the commit message with the cherry pick lines | |
| 17:38:12 | rloo | well, if i manually do that -- there won't be a conflict in the wallaby one (if i recall) cuz the change was similar to the xena one. but i didn't tell you that, i'll fix the commit messages... | |
| 17:39:06 | sean-k-mooney | right so i do not normllay remove the confit bit in that case although i know other do | |
| 17:39:16 | sean-k-mooney | i do if others ask | |
| 17:39:27 | rloo | (and if someone had time to fix that UI so it doesn't allow cherry picking from master to n-2+ stable branches, heh) | |
| 17:39:58 | sean-k-mooney | one thing i have not tested is if the behvior change if the patch is merged | |