| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-31 | |||
| 16:41:34 | bauzas | isn't it changing the logic ? | |
| 16:41:37 | gibi | based on the impl it is backportable | |
| 16:41:52 | bauzas | wait | |
| 16:42:09 | gibi | it has a dependency on https://review.opendev.org/q/topic:bug%252F1628606 which is being backported | |
| 16:42:14 | bauzas | this is post_live_mig_at_dest() | |
| 16:42:27 | bauzas | which is run on the dest | |
| 16:42:54 | bauzas | my question is, in a rolling upgrade scenario with operators moving workloads from old compute to new compute | |
| 16:43:08 | bauzas | would that break them ? | |
| 16:43:27 | gibi | bauzas: the bug is ther until the dest is upgraded | |
| 16:44:00 | bauzas | we say we gonna rely on https://review.opendev.org/c/openstack/nova/+/791135 | |
| 16:44:26 | bauzas | but correct me if I'm wrong, this won't happen for a A to B livemig if A is old, nope ? | |
| 16:44:56 | bauzas | this chit-chat limbo dance between two hosts is confusing | |
| 16:45:09 | bauzas | I never know which compute runs which codepath | |
| 16:45:11 | gibi | this patch trying to fix a bug that happens on the dest after the live migration finished | |
| 16:45:27 | gibi | at that point there is no return to the source host | |
| 16:45:56 | gibi | Amit fixed that in this case we set the instance to point to the dest host so a hard reboot can recover the instance | |
| 16:46:24 | bauzas | but we backported Amit's patch, right? | |
| 16:46:29 | gibi | yes | |
| 16:46:36 | bauzas | ok, so we're safe | |
| 16:46:46 | gibi | Uggla had a case where a very similar failure could leave the instance in Migarting state | |
| 16:47:00 | gibi | this fix is putting the instance in error in this case too | |
| 16:47:12 | bauzas | I see | |
| 16:47:35 | bauzas | ok, then to answer Uggla's question, I don't see any controversy to backport such change once we merge it | |
| 16:47:58 | gibi | yepp | |
| 16:48:24 | Uggla | down to ? | |
| 16:48:34 | gibi | the same version as Amit's | |
| 16:48:40 | bauzas | yup | |
| 16:48:48 | gibi | I think that is proposed back to train | |
| 16:48:49 | bauzas | which is train iirc | |
| 16:48:52 | bauzas | yup | |
| 16:49:04 | bauzas | but as we said, we're on hold on wallaby | |
| 16:49:13 | bauzas | nothing prevents us to do the work tho | |
| 16:50:23 | bauzas | can we end the meeting then ? | |
| 16:50:55 | Uggla | ok 4 me. | |
| 16:51:08 | gibi | nothing else from me | |
| 16:51:57 | bauzas | then | |
| 16:51:59 | bauzas | thanks all | |
| 16:52:04 | bauzas | #endmeeting | |
| 16:52:04 | opendevmeet | Meeting ended Tue Jan 31 16:52:04 2023 UTC. Information about MeetBot at http://wiki.debian.org/MeetBot . (v 0.1.4) | |
| 16:52:04 | opendevmeet | Minutes: https://meetings.opendev.org/meetings/nova/2023/nova.2023-01-31-16.00.html | |
| 16:52:04 | opendevmeet | Minutes (text): https://meetings.opendev.org/meetings/nova/2023/nova.2023-01-31-16.00.txt | |
| 16:52:04 | opendevmeet | Log: https://meetings.opendev.org/meetings/nova/2023/nova.2023-01-31-16.00.log.html | |
| 16:55:40 | elodilles | thanks o/ | |
| 16:56:15 | kashyap | sean-k-mooney: Running the stable/xena locally gives me this yak to shave - https://paste.opendev.org/show/bsRpOt3yk8k6LUw5D9P0/ | |
| 16:57:50 | elodilles | kashyap: maybe i'm wrong but this can be fixed by updating/installing explicitly setuptools? | |
| 16:58:07 | kashyap | Probably; I'm on Fedora 36 | |
| 16:58:25 | kashyap | elodilles: I'm trying to see if this erroneous failure of stable/xena related to my backport or not (it doesn't seem so) | |
| 16:58:28 | kashyap | https://zuul.opendev.org/t/openstack/build/796a6b05d72c4bbd87a3375028d43a1f | |
| 16:58:39 | kashyap | This one: nova.tests.unit.virt.libvirt.test_driver.LibvirtConnTestCase.test_check_can_live_migrate_dest_numa_lm [0.046959s] | |
| 16:58:43 | kashyap | (And that's the backport - https://review.opendev.org/c/openstack/nova/+/851205) | |
| 16:59:22 | sean-k-mooney | oh thats a known issue | |
| 16:59:51 | elodilles | yes, a known one and fixed in upstream gate i think | |
| 16:59:55 | sean-k-mooney | use_2to3 was remvoed in a setuptools verison | |
| 17:00:29 | kashyap | elodilles: Ah, thx. gibi also pointed that there's a rename, hence the fail: https://review.opendev.org/c/openstack/nova/+/871975 | |
| 17:00:32 | kashyap | Thanks, gibi! | |
| 17:00:35 | sean-k-mooney | https://github.com/gibizer/openstack-tox-docker/blob/main/ussuri/Dockerfile | |
| 17:00:38 | gibi | I thiunk the yoga container from here https://github.com/gibizer/openstack-tox-docker work on xena too | |
| 17:00:52 | sean-k-mooney | gibi: yes it does | |
| 17:01:10 | sean-k-mooney | kashyap: we swapped form the unmaintained suds_junko repo to a differnt one | |
| 17:01:26 | sean-k-mooney | but that is not your issue | |
| 17:01:53 | sean-k-mooney | you will need to clamp your pip/virtualevn and tox version | |
| 17:03:49 | elodilles | yepp. in upstream xena has newer versions (ubuntu) that's why we needed this up till ussuri ( https://review.opendev.org/c/openstack/nova/+/810461 ) | |
| 17:04:22 | elodilles | so i guess, in fedora this is needed in xena as well | |
| 17:16:15 | opendevreview | Sahid Orentino Ferdjaoui proposed openstack/nova master: api: extend evacuate instance to support target state https://review.opendev.org/c/openstack/nova/+/858384 | |
| 17:16:45 | sahid | addressed minor comments from Rajesh Tailor, thank you ! | |
| 18:32:32 | dansmith | sean-k-mooney: gibi bauzas: Things like this https://review.opendev.org/c/openstack/nova/+/872204 are incredibly difficult to make "work" in the functional tests because of all the ways and patterns we start and restart multiple fake computes | |
| 18:33:23 | bauzas | gdam shit | |
| 18:33:40 | dansmith | how terrible would it be if we just always mock out those host consistency checks in functionals and rely on unit testing to cover them, and the tempest jobs to really cover the regular happy path(s) ? | |
| 18:34:13 | dansmith | because the levels of stupid mocking to get even some of them to pass are probably worse and more complex than just not ever running those in the functionals | |
| 18:35:34 | bauzas | dansmith: I need to look at your failures but unfortunately I need to quit today (EOB) | |
| 18:36:28 | dansmith | bauzas: okay well, I'm just talking about a general read on the principle of the thing, but ... okay | |
| 18:37:27 | bauzas | dansmith: if you want, mock them indeed and just leave the unittests for checking them | |
| 18:38:24 | dansmith | ack | |
| 18:38:38 | dansmith | probably need a read from the other two before I go down that route | |
| 18:39:05 | sean-k-mooney | you could use a class constant to enable/disbale it | |
| 18:39:13 | sean-k-mooney | but enabel it by default in the base test | |
| 18:39:26 | dansmith | to enable/disable the mocking you mean? | |
| 18:39:32 | sean-k-mooney | yep | |
| 18:39:43 | sean-k-mooney | kind of like the microversion class constant | |
| 18:40:20 | sean-k-mooney | or the db one | |
| 18:40:45 | dansmith | okay, I'm not sure that will get me much, assuming you mean "so you can test it in one functional that does things in a specific way" because of the way the rest of the singleton virt node mocking works | |
| 18:40:53 | dansmith | but I can try | |
| 18:41:24 | sean-k-mooney | no i mean by default mock it out so it passes | |
| 18:41:42 | sean-k-mooney | and if you need to test it then you would set MOCK_STABLE_UUID=False | |
| 18:41:51 | sean-k-mooney | where you are explictly testing something that cares | |
| 18:42:22 | sean-k-mooney | following this pattern https://github.com/openstack/nova/blob/master/nova/test.py#L158-L173 | |
| 18:43:08 | dansmith | what I'm saying is, even with the mock disabled for one test, the virt node mock always returns None, so testing this in a full stack is kinda difficult without *also* making that parameterized | |
| 18:43:28 | sean-k-mooney | oh ok | |
| 18:43:39 | dansmith | but perhaps if I put them both behind that flag I'll get what I need, I'll have to see | |
| 18:44:10 | sean-k-mooney | this is just needed for the last patch correct? | |
| 18:44:28 | sean-k-mooney | as in you already worked though the other test isseus in the previous ones | |
| 18:45:27 | dansmith | yeah | |
| 18:46:01 | sean-k-mooney | so do you want to go up one level and just mock out _ensure_existing_node_identity and the check function for host and hypervisor hostname | |
| 18:46:53 | dansmith | that's what I was going to do, after two hours of trying to mock only the second call to it, | |
| 18:47:03 | dansmith | but there are just too many permutations of start, restart, stop, start, etc | |
| 18:47:36 | dansmith | but let me try to the flag for both the node mock and that one and see if i think it results in something meaningful | |
| 18:48:18 | dansmith | I feel like it will probably end up with 10,000 tests having that set, just so one can have it unset, but still need mocks to make it not very useful, vs. just unit testing it in isolation | |
| 18:48:23 | dansmith | but I'll see | |
| 18:49:06 | dansmith | tbh I wasn't thinking about tying the node mock to the compute _ensure one, so will try that first | |
| 19:05:34 | dansmith | actually, maybe this will be better anyway and then I can write some dedicated lifecycle tests to simulate the manual testing we've been doing with devstack | |