Earlier  
Posted Nick Remark
#openstack-nova - 2023-02-10
10:40:50 bauzas Ogame, got it
10:41:14 bauzas that one stole too much of my free time
10:49:25 bauzas gibi: about https://review.opendev.org/c/openstack/nova/+/821228/6/nova/virt/libvirt/host.py#745 I wonder whether we really need to call *again* getCPUMap()
10:49:43 bauzas gibi: the existing logic just considers to return the cpu blindless
10:49:48 bauzas blindly
10:50:27 bauzas but I see your point
10:50:32 bauzas nevermind
10:51:24 gibi bauzas: you probably still need to differentiate between non existent cpu ids in the dedicated_set and existing but offined cpus
10:52:00 bauzas yeah, I'm about adding a get_available_cpus() which will return all CPUs for the map
10:53:23 gibi yeah that will work
10:59:45 kashyap bauzas: What exactly does getCPUMap() fetch? The docs only say "Get node CPU information"
11:00:12 kashyap Ah, it maps to `virsh cpu-stats`
11:00:57 bauzas yup
11:01:01 kashyap Oh, interesting. When I run `virsh cpu-stats` on my Fedora 36 VM for a guest, it gives me:
11:01:05 bauzas https://libvirt.org/html/libvirt-libvirt-host.html#virNodeGetCPUMap
11:01:12 kashyap $> sudo virsh cpu-stats 1
11:01:13 kashyap error: Operation not supported: operation 'getCpuacctPercpuUsage' not supported for backend 'cgroup V2'
11:01:13 kashyap error: Failed to retrieve CPU statistics for domain 'el8-vm1'
11:01:25 bauzas [sbauza@sbauza temp]$ python
11:01:26 kashyap bauzas: Yeah, was reading. It looks like there's some accounting to be done w.r.t CGroups version
11:01:26 bauzas >>> conn = libvirt.open('qemu:///system')
11:01:26 bauzas >>> import libvirt
11:01:26 bauzas Type "help", "copyright", "credits" or "license" for more information.
11:01:26 bauzas Python 3.11.1 (main, Jan 6 2023, 00:00:00) [GCC 12.2.1 20221121 (Red Hat 12.2.1-4)] on linux
11:01:27 bauzas >>> conn.getCPUMap()
11:01:29 bauzas (8, [True, True, True, True, True, True, True, True], 8)
11:01:55 kashyap Yep
11:10:42 sean-k-mooney you can also just get the online cores form sysfs
11:11:18 sean-k-mooney if this is an issue but im fine with using libvirt's api for this
11:16:12 bauzas I'm done with the new rev, just updating the upper patch now
11:24:22 sean-k-mooney my steamdeck is charging so im currently still at my work laptop.(playing factorio) so gibi if your happy to review bauzas serise and it looks good to you ping me when your done and i can then do a final pass over it quickly and we can likely merge that today. with that said i have a doctors apointmen in a littel over 3 hours so ill be away after that.
11:24:38 bauzas that's a love
11:24:52 auniyal O/
11:24:57 gibi sean-k-mooney: ack
11:25:04 auniyal in devstack is there a way to see nova-manage logs
11:25:20 bauzas I'm honestly torn, I don't know whether we should really pay attention to ignoring errors
11:25:36 bauzas and whether we should be cautious
11:25:55 bauzas honestly, we create a temp dir, so I don't except problems besides the full disk problem
11:26:11 sean-k-mooney you should not get any now with the way the fixture works
11:26:17 bauzas so I'll turn into using the .cleanup() method
11:26:22 bauzas and meh
11:26:30 bauzas meh to ignoring errors
11:26:42 gibi I'm fine ignoring errors
11:26:49 gibi during delete of a temp dir
11:27:00 gibi we will never reuse the temp dir as it has a random postfix
11:27:05 sean-k-mooney the flag is only there in 3.10
11:27:09 gibi if my /tmp fills up that is on me
11:27:10 sean-k-mooney and we need to support 3.8
11:27:17 gibi aah
11:27:21 bauzas gibi: that's the problem
11:27:26 sean-k-mooney so its fine for it to error
11:27:28 sean-k-mooney it wont
11:27:29 gibi ahh
11:27:30 bauzas we can't say 'ignore errors' if we use cleanup
11:27:38 bauzas hence me torn
11:27:41 gibi I'm fine both ways
11:27:41 sean-k-mooney i could have errored when we had the copy of sysfs because of some permision
11:27:51 sean-k-mooney but now its just normal files owned by us
11:27:54 bauzas correct
11:28:00 gibi if it starts failing on cleanup then we will switch to shutil
11:28:01 sean-k-mooney so it wont error unless there is a disk issue which si out of scope
11:28:08 bauzas yup, this ^
11:28:09 sean-k-mooney +1
11:28:44 bauzas I'll add a comment explaining the risk and how to mitigate it if we see it in CI
11:29:39 gibi cool
11:39:48 bauzas sean-k-mooney: gibi: I haven't yet written the docs patch as it requires a bit of effort, so I'll keep your comments on it unresolved but don't misunderstand me, surely I'll do it in a subsequent patch once I'm done (monday morning hopefully)
11:39:58 bauzas I just wanna give chance to other people to get reviewed
11:40:14 gibi bauzas: sure, doc is OK after FF
11:40:22 sean-k-mooney ya it can be a seperate patch
11:40:47 bauzas the core #0 note is actually very important
11:40:49 bauzas TIL about it
11:41:13 bauzas but yeah that makes sense from an OS perspective
11:41:21 bauzas you always rely on that core to be available
11:41:50 bauzas that's the most portable assumption
11:42:04 sean-k-mooney ya it generally need one core that can always be used to handel interupts and you know turn on the others
11:43:14 sean-k-mooney bauzas: that is why there is no online file in /sys/bus/cpu/devices/cpu0/
11:43:24 sean-k-mooney but there is in all the rest
11:44:07 bauzas ooooooooh
11:44:39 bauzas but you can write an online file and set 0 into it for core #0, right ?
11:45:13 bauzas that's quite a destructive CPU equivalent of disk's rm -rf /
11:45:25 bauzas except it's stateless
11:49:40 sean-k-mooney it will be ignored
11:49:55 sean-k-mooney you can create that file but it wont do anything as far as i am aware
11:50:52 sean-k-mooney by the way https://github.com/SeanMooney/arbiterd/blob/master/src/arbiterd/common/cpu.py#L58-L62
11:51:02 sean-k-mooney is how i got the aviable cpus
11:52:02 sean-k-mooney gibi: bauzas actully also for context
11:52:04 sean-k-mooney https://github.com/SeanMooney/arbiterd/blob/master/src/arbiterd/common/cpu.py#L106-L107
11:52:14 sean-k-mooney that default 1 was because of this
11:52:19 bauzas gtk
11:52:24 sean-k-mooney https://github.com/SeanMooney/arbiterd/blob/master/src/arbiterd/common/cpu.py#L106-L114
11:52:35 sean-k-mooney thats also why i did the get_online check in set online
11:52:46 sean-k-mooney to not do the write
11:52:59 sean-k-mooney i proably should have left a code comment for that...
11:53:26 sean-k-mooney ye removed that optimisation but that was actully the real reason i did that in the poc. it was not an optimisation
11:53:37 sean-k-mooney it was to workaround the cpu0 weridness
11:54:28 sean-k-mooney although to be fair i didnt fix that on the set offline path
11:54:30 sean-k-mooney so meh
12:04:01 opendevreview Sylvain Bauza proposed openstack/nova master: libvirt: let CPUs be power managed https://review.opendev.org/c/openstack/nova/+/821228
12:04:02 opendevreview Sylvain Bauza proposed openstack/nova master: Enable cpus when an instance is spawning https://review.opendev.org/c/openstack/nova/+/868237
12:04:06 bauzas gibi: sean-k-mooney: ^

Earlier   Later