| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-13 | |||
| 21:41:41 | sean-k-mooney[m] | is that how we have the functional job configured in nova/osc-placement | |
| 21:41:57 | gmann | if we change it to test with master only then we should have another job to test with the released version | |
| 21:42:20 | sean-k-mooney[m] | well its testing with master only now | |
| 21:42:24 | gmann | as comments says, we can test the master placement by replacing the deps line | |
| 21:42:38 | gmann | yes, in osc-placement and released version in nova | |
| 21:42:47 | sean-k-mooney[m] | so your propeosing change to using the released version | |
| 21:42:56 | sean-k-mooney[m] | yep | |
| 21:43:07 | sean-k-mooney[m] | so it would be nice for depends-on to work | |
| 21:43:19 | sean-k-mooney[m] | and in generall im fine with used the released verison | |
| 21:43:40 | sean-k-mooney[m] | but im questioning if the tox job in osc-placment will give the depends on behavior today | |
| 21:43:58 | gmann | While doing it for stable branch and I checked how Nova does I thought of doing the same for osc-placement also | |
| 21:45:15 | sean-k-mooney[m] | ok so we have the required projec tin the zuul.yaml | |
| 21:45:16 | sean-k-mooney[m] | https://github.com/openstack/osc-placement/blob/master/.zuul.yaml | |
| 21:45:38 | gmann | yeah | |
| 21:45:45 | sean-k-mooney[m] | yep i know i was just checking if we were using the jobs from the default template or if we had already overriden it | |
| 21:46:00 | sean-k-mooney[m] | since i did not see that in the patch you linked | |
| 21:46:11 | sean-k-mooney[m] | i was expecting both in the same patch | |
| 21:47:58 | sean-k-mooney[m] | ok this should be fine as is | |
| 21:48:59 | gmann | ack | |
| 21:52:55 | clarkb | I'm not sure the comment is correct though since siblings will be used | |
| 21:53:10 | clarkb | it will always be the latest version of the branch in the gate not the release | |
| 23:22:37 | dansmith | cripes, we're never going to get the rpc spam thing landed | |
| 23:22:44 | dansmith | seen this a few times now as well: https://zuul.opendev.org/t/openstack/build/f5aa5edd4d354c2685fc1f3e13d0ef77 | |
| 23:22:59 | dansmith | saying one of the tempest workers crashed | |
| 23:23:06 | dansmith | seems unlikely to me | |
| #openstack-nova - 2023-01-14 | |||
| 01:55:42 | opendevreview | Merged openstack/nova master: libvirt: Add encryption support to qemu-img create command https://review.opendev.org/c/openstack/nova/+/826752 | |
| 01:55:50 | opendevreview | Merged openstack/nova master: libvirt: Report ephemeral encryption traits based on imagebackend https://review.opendev.org/c/openstack/nova/+/826753 | |
| 13:11:02 | opendevreview | Artom Lifshitz proposed openstack/nova master: Microversion 2.94: FQDN in hostname https://review.opendev.org/c/openstack/nova/+/869812 | |
| 13:18:54 | opendevreview | Artom Lifshitz proposed openstack/nova master: Microversion 2.94: FQDN in hostname https://review.opendev.org/c/openstack/nova/+/869812 | |
| #openstack-nova - 2023-01-16 | |||
| 07:39:50 | opendevreview | Merged openstack/nova stable/ussuri: [compute] always set instance.host in post_livemigration https://review.opendev.org/c/openstack/nova/+/864007 | |
| 07:42:23 | opendevreview | Amit Uniyal proposed openstack/nova stable/train: Adds a repoducer for post live migration fail https://review.opendev.org/c/openstack/nova/+/863806 | |
| 07:42:24 | opendevreview | Amit Uniyal proposed openstack/nova stable/train: [compute] always set instance.host in post_livemigration https://review.opendev.org/c/openstack/nova/+/864055 | |
| 08:48:10 | gibi | dansmith: I saw such interpreter crashes before. It is really a segfault of the python interpreter based on dmesg. Unfortunately it happens randomly afais. | |
| 08:53:47 | gibi | dansmith: hm, but this time it was OOM | |
| 09:16:02 | viks__ | hi, is there a way to set say `/var/lib/nova1` instead of `/var/lib/nova` ? i could not find any configuration to do that in `nova.conf` | |
| 09:43:08 | gibi | viks__: https://docs.openstack.org/nova/latest/configuration/config.html#DEFAULT.instances_path and https://docs.openstack.org/nova/latest/configuration/config.html#DEFAULT.state_path are the way I think | |
| 09:48:43 | viks__ | gibi: thanks... got it.. actually i was searching for `/var/lib/nova` in sample conf so i did not find it before... anyway can i set multiple values for it say for eg: `/var/lib/nova,/var/lib/nova1` where they are differnt 2 mount points ? | |
| 09:50:54 | gibi | viks__: you can only set a single path | |
| 09:52:48 | viks__ | gibi: ok.. thanks.. one more thing.. when to use `instances_path` if `state_path` itself will do the job? any suggestions? | |
| 09:55:18 | gibi | viks__: if you want to store the instance local disks in a different place then the nova lock files then instances_path will let you separate the instance disk from the lock files | |
| 10:04:32 | viks__ | gibi: ok.. thanks | |
| 10:14:46 | opendevreview | Sahid Orentino Ferdjaoui proposed openstack/nova master: api: extend evacuate instance to support target state https://review.opendev.org/c/openstack/nova/+/858384 | |
| 10:14:46 | opendevreview | Sahid Orentino Ferdjaoui proposed openstack/nova master: compute: enhance compute evacuate instance to support target state https://review.opendev.org/c/openstack/nova/+/858383 | |
| 10:32:43 | gibi | dansmith: opened the bug for the OOM https://bugs.launchpad.net/nova/+bug/2002951 | |
| 10:54:17 | sean-k-mooney | viks__: if you cant use a single mount path for soem reason you could fake it using lvm voluems to combine multipel disk into one or do it at thte file system level instead of block level using mergerfs https://manpages.ubuntu.com/manpages/impish/man1/mergerfs.1.html | |
| 10:55:38 | sean-k-mooney | nova need everything to be in a single directory on the file system but we dont really care where that folder comes form or how you created it | |
| 10:56:31 | bauzas | auniyal: so, about what we discussed for https://bugs.launchpad.net/nova/+bug/1996732 | |
| 10:57:11 | bauzas | auniyal: what you need to do first is to check how to look at the host_state.failed_builds value | |
| 10:58:16 | sean-k-mooney | all you need to do is add a new exception that inherits form the existign one and then where we incremente the value skip it if its the new excpetion | |
| 10:58:36 | sean-k-mooney | the late affinity failure will use the new expction and not be counted | |
| 10:58:54 | sean-k-mooney | but because it inherits form the orginal any clean up that was previousl done will still be done | |
| 10:59:36 | sean-k-mooney | so ya one of the first steps is find out where we modify the build failure value | |
| 10:59:51 | sean-k-mooney | and also where we raise the current excetption for the affinity check | |
| 11:00:29 | sean-k-mooney | after that you can add the new expction and modify both to use it | |
| 11:00:34 | bauzas | sean-k-mooney: the concern of auniyal was how to test iut | |
| 11:04:06 | sean-k-mooney | out side of functional tests it would be tricky to do end to end but unit test for the excption raising and functional test for the end to end interaction. there is no point doing tepest type testing since fault inject will be needed and that is not somethign tempest is good for | |
| 11:06:07 | bauzas | sean-k-mooney: I think it's possible to have a functest for it | |
| 11:06:19 | bauzas | but we need to be able to verify the stats | |
| 11:06:27 | viks__ | sean-k-mooney: thanks for the suggestions | |
| 11:09:29 | bauzas | sean-k-mooney: if we have a functest that creates a group with anti-affinity policy without having the filter, then we can create two instances asking for the same host | |
| 11:09:34 | auniyal | sean-k-mooney, bauzas, suppose we have only one host hostA, an instance with server group having a anti-affinity is present hostA, so if user create anothere instance of same group, that time it will go for reschedule, and hence increase the counter of build faild | |
| 11:09:40 | auniyal | is this a valid scenario | |
| 11:09:44 | bauzas | sean-k-mooney: and then we could verify the stats value | |
| 11:09:44 | auniyal | to test | |
| 11:10:20 | bauzas | auniyal: yeah | |
| 11:10:31 | bauzas | auniyal: but then you need to verify the counter | |
| 11:10:46 | bauzas | https://github.com/openstack/nova/blob/2eb358cdcec36fcfe5388ce6982d2961ca949d0a/nova/compute/resource_tracker.py#L1978-L1980 | |
| 11:11:20 | bauzas | which is a defaultdict set here https://github.com/openstack/nova/blob/2eb358cdcec36fcfe5388ce6982d2961ca949d0a/nova/compute/resource_tracker.py#L99 | |
| 11:12:08 | bauzas | so the stats is updated here https://github.com/openstack/nova/blob/2eb358cdcec36fcfe5388ce6982d2961ca949d0a/nova/compute/stats.py#L143 | |
| 11:12:47 | bauzas | so, as you'll see, this is a keyed dict | |
| 11:13:08 | bauzas | with a functional test, you can introspect that compute.stats dict | |
| 11:13:47 | bauzas | and see that the value of that dict for the key 'failed_builds' is incremented by 1 | |
| 11:13:57 | bauzas | after creating inst2 | |
| 11:14:09 | bauzas | and that's what we want to change | |
| 11:14:26 | bauzas | auniyal: ^ | |
| 11:16:32 | auniyal | ack, this will tell us isntance failed again and agian | |
| 11:19:17 | gibi | auniyal: no, if you have one host with an instance in an anti-affinity group and you try to schedule the second instance to the same group then it will not trigger a reschedule | |
| 11:19:36 | gibi | it will simply fail the scheduling with NoValidHost | |
| 11:19:45 | auniyal | yes | |
| 11:20:07 | bauzas | yes you need a test with 2 nodes, don't disagree | |
| 11:20:26 | gibi | bauzas: two nodes will not help either as the scheduler will pick the other host | |
| 11:20:35 | bauzas | gibi: not if you trick it :) | |
| 11:20:38 | gibi | to trigger a late affinity check failure (and hence the build failure counter increase) you need to have two parallel scheduling request | |
| 11:21:10 | bauzas | gibi: my proposal is simplier with a functest, we have code snippets for tricking the scheduler | |
| 11:22:01 | bauzas | or you could use the az hack | |
| 11:22:08 | bauzas | it will skip the scheduler | |
| 11:23:18 | gibi | (alternatively we could try to remove the anti-affinity filter from the config then the scheduler will allow both VMs to the same host, and the second will fail the late affinity check there) | |
| 11:26:58 | bauzas | gibi: ah, you missed then my point | |
| 11:27:06 | bauzas | (12:09:29) bauzas: sean-k-mooney: if we have a functest that creates a group with anti-affinity policy without having the filter, then we can create two instances asking for the same host | |
| 11:27:30 | bauzas | we indeed need a functest that *doesn't* use the AntiAffinityFilter | |
| 11:27:48 | bauzas | faking the scheduler is just for making sure we land instances on the same host | |
| 11:28:55 | gibi | bauzas: ack, then we thought about the same thing. cool :) | |
| 11:29:17 | auniyal | can we reproduce it manually | |
| 11:31:28 | auniyal | I understand reschdule is correct but its should be counted | |
| 11:32:04 | auniyal | so we just need to verify this before sauing buld failed at https://github.com/openstack/nova/blob/2eb358cdcec36fcfe5388ce6982d2961ca949d0a/nova/compute/manager.py#L2265 | |
| 11:32:33 | auniyal | *it should NOT be counted | |
| 11:37:10 | bauzas | auniyal: you can reproduce it with devstack | |
| 11:37:19 | bauzas | auniyal: make sure the filter is disabled | |
| 11:37:34 | bauzas | and force to create two instances with the same group on the same host | |
| 11:37:43 | bauzas | that should work | |