Earlier  
Posted Nick Remark
#openstack-nova - 2023-01-13
21:46:00 sean-k-mooney[m] since i did not see that in the patch you linked
21:46:11 sean-k-mooney[m] i was expecting both in the same patch
21:47:58 sean-k-mooney[m] ok this should be fine as is
21:48:59 gmann ack
21:52:55 clarkb I'm not sure the comment is correct though since siblings will be used
21:53:10 clarkb it will always be the latest version of the branch in the gate not the release
23:22:37 dansmith cripes, we're never going to get the rpc spam thing landed
23:22:44 dansmith seen this a few times now as well: https://zuul.opendev.org/t/openstack/build/f5aa5edd4d354c2685fc1f3e13d0ef77
23:22:59 dansmith saying one of the tempest workers crashed
23:23:06 dansmith seems unlikely to me
#openstack-nova - 2023-01-14
01:55:42 opendevreview Merged openstack/nova master: libvirt: Add encryption support to qemu-img create command https://review.opendev.org/c/openstack/nova/+/826752
01:55:50 opendevreview Merged openstack/nova master: libvirt: Report ephemeral encryption traits based on imagebackend https://review.opendev.org/c/openstack/nova/+/826753
13:11:02 opendevreview Artom Lifshitz proposed openstack/nova master: Microversion 2.94: FQDN in hostname https://review.opendev.org/c/openstack/nova/+/869812
13:18:54 opendevreview Artom Lifshitz proposed openstack/nova master: Microversion 2.94: FQDN in hostname https://review.opendev.org/c/openstack/nova/+/869812
#openstack-nova - 2023-01-16
07:39:50 opendevreview Merged openstack/nova stable/ussuri: [compute] always set instance.host in post_livemigration https://review.opendev.org/c/openstack/nova/+/864007
07:42:23 opendevreview Amit Uniyal proposed openstack/nova stable/train: Adds a repoducer for post live migration fail https://review.opendev.org/c/openstack/nova/+/863806
07:42:24 opendevreview Amit Uniyal proposed openstack/nova stable/train: [compute] always set instance.host in post_livemigration https://review.opendev.org/c/openstack/nova/+/864055
08:48:10 gibi dansmith: I saw such interpreter crashes before. It is really a segfault of the python interpreter based on dmesg. Unfortunately it happens randomly afais.
08:53:47 gibi dansmith: hm, but this time it was OOM
09:16:02 viks__ hi, is there a way to set say `/var/lib/nova1` instead of `/var/lib/nova` ? i could not find any configuration to do that in `nova.conf`
09:43:08 gibi viks__: https://docs.openstack.org/nova/latest/configuration/config.html#DEFAULT.instances_path and https://docs.openstack.org/nova/latest/configuration/config.html#DEFAULT.state_path are the way I think
09:48:43 viks__ gibi: thanks... got it.. actually i was searching for `/var/lib/nova` in sample conf so i did not find it before... anyway can i set multiple values for it say for eg: `/var/lib/nova,/var/lib/nova1` where they are differnt 2 mount points ?
09:50:54 gibi viks__: you can only set a single path
09:52:48 viks__ gibi: ok.. thanks.. one more thing.. when to use `instances_path` if `state_path` itself will do the job? any suggestions?
09:55:18 gibi viks__: if you want to store the instance local disks in a different place then the nova lock files then instances_path will let you separate the instance disk from the lock files
10:04:32 viks__ gibi: ok.. thanks
10:14:46 opendevreview Sahid Orentino Ferdjaoui proposed openstack/nova master: api: extend evacuate instance to support target state https://review.opendev.org/c/openstack/nova/+/858384
10:14:46 opendevreview Sahid Orentino Ferdjaoui proposed openstack/nova master: compute: enhance compute evacuate instance to support target state https://review.opendev.org/c/openstack/nova/+/858383
10:32:43 gibi dansmith: opened the bug for the OOM https://bugs.launchpad.net/nova/+bug/2002951
10:54:17 sean-k-mooney viks__: if you cant use a single mount path for soem reason you could fake it using lvm voluems to combine multipel disk into one or do it at thte file system level instead of block level using mergerfs https://manpages.ubuntu.com/manpages/impish/man1/mergerfs.1.html
10:55:38 sean-k-mooney nova need everything to be in a single directory on the file system but we dont really care where that folder comes form or how you created it
10:56:31 bauzas auniyal: so, about what we discussed for https://bugs.launchpad.net/nova/+bug/1996732
10:57:11 bauzas auniyal: what you need to do first is to check how to look at the host_state.failed_builds value
10:58:16 sean-k-mooney all you need to do is add a new exception that inherits form the existign one and then where we incremente the value skip it if its the new excpetion
10:58:36 sean-k-mooney the late affinity failure will use the new expction and not be counted
10:58:54 sean-k-mooney but because it inherits form the orginal any clean up that was previousl done will still be done
10:59:36 sean-k-mooney so ya one of the first steps is find out where we modify the build failure value
10:59:51 sean-k-mooney and also where we raise the current excetption for the affinity check
11:00:29 sean-k-mooney after that you can add the new expction and modify both to use it
11:00:34 bauzas sean-k-mooney: the concern of auniyal was how to test iut
11:04:06 sean-k-mooney out side of functional tests it would be tricky to do end to end but unit test for the excption raising and functional test for the end to end interaction. there is no point doing tepest type testing since fault inject will be needed and that is not somethign tempest is good for
11:06:07 bauzas sean-k-mooney: I think it's possible to have a functest for it
11:06:19 bauzas but we need to be able to verify the stats
11:06:27 viks__ sean-k-mooney: thanks for the suggestions
11:09:29 bauzas sean-k-mooney: if we have a functest that creates a group with anti-affinity policy without having the filter, then we can create two instances asking for the same host
11:09:34 auniyal sean-k-mooney, bauzas, suppose we have only one host hostA, an instance with server group having a anti-affinity is present hostA, so if user create anothere instance of same group, that time it will go for reschedule, and hence increase the counter of build faild
11:09:40 auniyal is this a valid scenario
11:09:44 bauzas sean-k-mooney: and then we could verify the stats value
11:09:44 auniyal to test
11:10:20 bauzas auniyal: yeah
11:10:31 bauzas auniyal: but then you need to verify the counter
11:10:46 bauzas https://github.com/openstack/nova/blob/2eb358cdcec36fcfe5388ce6982d2961ca949d0a/nova/compute/resource_tracker.py#L1978-L1980
11:11:20 bauzas which is a defaultdict set here https://github.com/openstack/nova/blob/2eb358cdcec36fcfe5388ce6982d2961ca949d0a/nova/compute/resource_tracker.py#L99
11:12:08 bauzas so the stats is updated here https://github.com/openstack/nova/blob/2eb358cdcec36fcfe5388ce6982d2961ca949d0a/nova/compute/stats.py#L143
11:12:47 bauzas so, as you'll see, this is a keyed dict
11:13:08 bauzas with a functional test, you can introspect that compute.stats dict
11:13:47 bauzas and see that the value of that dict for the key 'failed_builds' is incremented by 1
11:13:57 bauzas after creating inst2
11:14:09 bauzas and that's what we want to change
11:14:26 bauzas auniyal: ^
11:16:32 auniyal ack, this will tell us isntance failed again and agian
11:19:17 gibi auniyal: no, if you have one host with an instance in an anti-affinity group and you try to schedule the second instance to the same group then it will not trigger a reschedule
11:19:36 gibi it will simply fail the scheduling with NoValidHost
11:19:45 auniyal yes
11:20:07 bauzas yes you need a test with 2 nodes, don't disagree
11:20:26 gibi bauzas: two nodes will not help either as the scheduler will pick the other host
11:20:35 bauzas gibi: not if you trick it :)
11:20:38 gibi to trigger a late affinity check failure (and hence the build failure counter increase) you need to have two parallel scheduling request
11:21:10 bauzas gibi: my proposal is simplier with a functest, we have code snippets for tricking the scheduler
11:22:01 bauzas or you could use the az hack
11:22:08 bauzas it will skip the scheduler
11:23:18 gibi (alternatively we could try to remove the anti-affinity filter from the config then the scheduler will allow both VMs to the same host, and the second will fail the late affinity check there)
11:26:58 bauzas gibi: ah, you missed then my point
11:27:06 bauzas (12:09:29) bauzas: sean-k-mooney: if we have a functest that creates a group with anti-affinity policy without having the filter, then we can create two instances asking for the same host
11:27:30 bauzas we indeed need a functest that *doesn't* use the AntiAffinityFilter
11:27:48 bauzas faking the scheduler is just for making sure we land instances on the same host
11:28:55 gibi bauzas: ack, then we thought about the same thing. cool :)
11:29:17 auniyal can we reproduce it manually
11:31:28 auniyal I understand reschdule is correct but its should be counted
11:32:04 auniyal so we just need to verify this before sauing buld failed at https://github.com/openstack/nova/blob/2eb358cdcec36fcfe5388ce6982d2961ca949d0a/nova/compute/manager.py#L2265
11:32:33 auniyal *it should NOT be counted
11:37:10 bauzas auniyal: you can reproduce it with devstack
11:37:19 bauzas auniyal: make sure the filter is disabled
11:37:34 bauzas and force to create two instances with the same group on the same host
11:37:43 bauzas that should work
11:37:58 sean-k-mooney bauzas: you can contol what filters are used in the fucntest so that should not be a problem
11:38:07 bauzas introspecting the stats field would be a bit trickier but I think we log the stats with the DEBUG level
11:38:32 sean-k-mooney ya we likely do but if you really need to you can jsut go driect to the db
11:38:44 bauzas sean-k-mooney: yeah, and I even think we have a funtest fake filter for ensuring all instances go the same host
11:38:54 sean-k-mooney often w will use a spy function to intercept and recored such thigns
11:39:07 sean-k-mooney bauzas: we do yes
11:39:08 bauzas anyway, /me goes off for lunch
11:42:40 kashyap Hmm, on this bz: https://bugzilla.redhat.com/show_bug.cgi?id=2138381 (on CPU compatibility). A Red Hat customer-facing person says using the new CPU API still throws the same error. Actually removing the check is what works correctly:
11:42:44 kashyap https://review.opendev.org/c/openstack/nova/+/869587 -- libvirt: Remove compareCPU() check in _check_cpu_compatibility()
11:43:48 kashyap About debuggability concerns (if you remove the compareCPU() check: the same guy confirms you'd get the same error from libvirt. And it is still debuggable (which is what I said before)
11:44:33 sean-k-mooney kashyap: its much much much less debugable as you now need to boot a vm to triger it
11:46:16 sean-k-mooney since you have https://review.opendev.org/c/openstack/nova/+/869950 i think i would prefer if you abandoned https://review.opendev.org/c/openstack/nova/+/869587 and we proceded with the replacment instead
11:47:26 kashyap sean-k-mooney: Sure, I would also prefer the replacement
11:47:41 kashyap We can agree to disagree on "much much much"
11:48:05 kashyap I don't have the energy to argue much anyway; /me is still recovering from a bike accident

Earlier   Later