| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-06-02 | |||
| 15:27:47 | dansmith | kashyap: if it's just something we run post-crash we can add that as a one-off | |
| 15:27:56 | slaweq | kashyap sure, I'm on call now too | |
| 15:28:40 | kashyap | dansmith: What would be good is to capture both: `coredumpctl list | grep qemu`, and then for each QEMU PID log `coredumpctl info $PID` (I know ... I'm asking too much) | |
| 15:29:09 | kashyap | dansmith: E.g. see at the bottom here for the example output of `coredumpctl info $PID` - https://www.freedesktop.org/software/systemd/man/coredumpctl.html | |
| 15:29:59 | kashyap | The reason I ask is, I've successfully found several root-cause stack trace from it in the past. | |
| 15:31:02 | kashyap | dansmith: slaweq: Ah, scratch the above, we could even just get this post-crash: `coredumpctl -o qemu.coredump dump /usr/bin/qemu-system-x86_64` | |
| 15:32:17 | kashyap | slaweq: A quick question: the instance simply crashes when launching it? | |
| 15:32:51 | slaweq | kashyap I think it crashed during snapshoting | |
| 15:32:55 | slaweq | it was spawned properly | |
| 15:33:00 | opendevreview | Alexey Stupnikov proposed openstack/nova master: Optimize _local_delete calls by compute unit tests https://review.opendev.org/c/openstack/nova/+/844285 | |
| 15:33:51 | dansmith | kashyap: ah, running it on a specific pid would be much harder | |
| 15:34:26 | dansmith | running it like you describe is something we could hack in as a post job | |
| 15:34:32 | kashyap | dansmith: Nah, we can disregad the per-PID thing | |
| 15:34:47 | kashyap | Yeah, binary is easier indeed | |
| 15:35:23 | dansmith | okay after call(s) I can help hack that in if we need, but if it's pretty repeatable it might be easier to just try to repro locally | |
| 15:36:01 | kashyap | dansmith: Yeah, that's the next thing I'm looking. It looks like it's not any Ceph-based, and just plain local storage, IIRC | |
| 15:36:08 | dansmith | cool | |
| 15:37:49 | kashyap | slaweq: I'm just trying to find the precise test trigger. From looking at the 'n-cpu' log, snapshots seem to happen just fine. I'll look more after I'm done w/ this call | |
| 15:38:27 | slaweq | kashyap but IIUC nova logs, instance is gone during snapshoting process | |
| 15:38:37 | slaweq | please take Your time, it's not urgent for us for sure | |
| 15:38:52 | dansmith | kashyap: it's test_create_backup | |
| 15:38:55 | dansmith | https://storage.bhs.cloud.ovh.net/v1/AUTH_dcaab5e32b234d56b626f72581e3644c/zuul_opendev_logs_4a7/periodic/opendev.org/openstack/neutron/master/neutron-ovn-tempest-ovs-master-fedora/4a7f284/testr_results.html | |
| 15:39:25 | kashyap | (Yeah, just found it; thx) | |
| 15:43:08 | kashyap | dansmith: Thanks; so the above command I noted above "coredumpctl dump" will dump the most recent core dump. I guess we have to redirect it to a file | |
| 15:43:11 | opendevreview | Merged openstack/osc-placement master: Add Python3 zed unit tests https://review.opendev.org/c/openstack/osc-placement/+/835369 | |
| 15:43:39 | kashyap | dansmith: Err, ignore the above comment; the "-o" is the file. | |
| 15:43:58 | dansmith | kashyap: ack, are you going to try to repro locally first? | |
| 15:44:33 | kashyap | dansmith: On F35, just run the Tempest test, or construct a manual libvirt-based repro? | |
| 15:44:37 | dansmith | if it doesn't repro locally that might be interesting to know as well, like whether it's related to having run a lot of tests first, or that one thing always fails in isolation | |
| 15:44:55 | dansmith | kashyap: I just meant devstack, run that one tempest test | |
| 15:45:33 | kashyap | Ah; nod. I can't today; but I can give it a go tomorrow. | |
| 15:45:37 | dansmith | when I'm done here I can work on adding that as a post task, it just might take a bunch of iterations to get it right (based on experience) | |
| 15:45:45 | dansmith | okay, I'll give it a shot at least when I'm done here | |
| 15:47:00 | kashyap | dansmith: When you say "adding that as a post task" -- I take it you mean adding the above "coredumpctl ... dump", yeah? | |
| 15:47:08 | dansmith | yep | |
| 15:49:56 | kashyap | I have a deja vu about this test_create_backup test, reading its code | |
| 15:56:14 | kashyap | slaweq: When you get a minute, can you please file an upstream LP bug to track this? So we can keep all the investigation in one place? | |
| 16:01:12 | dansmith | kashyap: against what, nova? | |
| 16:01:26 | kashyap | Yeah, I'd say so | |
| 16:01:33 | kashyap | dansmith: I just looked at the compressed libvirtd.log | |
| 16:01:46 | kashyap | 2022-06-01 03:35:33.685+0000: 87576: error : qemuMonitorJSONCheckErrorFull:412 : internal error: unable to execute QEMU command 'blockdev-del': Failed to find node with node-name='libvirt-5-storage' | |
| 16:01:46 | kashyap | And I see a familiar libvirt error: | |
| 16:02:30 | kashyap | dansmith: In the past we found the same error earlier this year, and I recall working w/ libvirt folks to get a fix. But that was in a different context: https://listman.redhat.com/archives/libvir-list/2022-February/msg00790.html | |
| 16:03:11 | kashyap | dansmith: slaweq: The root TripleO (actually should've filed for Nova) bug was this where did the analysis: https://bugs.launchpad.net/tripleo/+bug/1959014 | |
| 16:03:26 | kashyap | If you open that last link, scroll from bottom for more signal | |
| 16:54:24 | dansmith | kashyap: yeah I remember that one.. so to be clear, you expect this is a different issue right? | |
| 17:09:24 | dansmith | kashyap: this is running, we'll see: https://review.opendev.org/c/openstack/devstack/+/844503 | |
| 17:10:56 | ricolin | bauzas: I think https://review.opendev.org/c/openstack/nova/+/830646 is ready for review now, could you kindly remove the -2 | |
| 17:11:54 | bauzas | ricolin: sure, lemme look | |
| 17:12:13 | ricolin | bauzas: thanks:) | |
| 17:12:33 | bauzas | ricolin: oh, yeah you created the bp and the spec, ta | |
| 17:12:57 | ricolin | bauzas: yeah, the spec merged:) | |
| 17:13:24 | ricolin | and I updated the implement patch accordingly | |
| 17:13:32 | ricolin | I think:) | |
| 17:20:05 | opendevreview | Rico Lin proposed openstack/nova master: Add traits for viommu model https://review.opendev.org/c/openstack/nova/+/844507 | |
| 17:21:17 | opendevreview | Artom Lifshitz proposed openstack/nova stable/ussuri: fake: Ensure need_legacy_block_device_info returns False https://review.opendev.org/c/openstack/nova/+/843950 | |
| 17:21:18 | opendevreview | Artom Lifshitz proposed openstack/nova stable/ussuri: Add a regression test for bug 1939545 https://review.opendev.org/c/openstack/nova/+/843951 | |
| 17:21:19 | opendevreview | Artom Lifshitz proposed openstack/nova stable/ussuri: compute: Ensure updates to bdms during pre_live_migration are saved https://review.opendev.org/c/openstack/nova/+/843952 | |
| 17:24:03 | ricolin | sean-k-mooney: this should be the last piece for libvirt-viommu-device implementation, but as I'm not familiar with traits, can you take a review on it and let me know if I do it right/wrong | |
| 17:24:05 | ricolin | https://review.opendev.org/c/openstack/nova/+/844507 | |
| 17:41:15 | sean-k-mooney | sure | |
| 17:41:30 | sean-k-mooney | just so you are aware the unit test will fail untill the trait is merged and released | |
| 17:41:50 | sean-k-mooney | but the tempest test shoudl be able to pass becauses depens on works for devstack jobs | |
| 17:41:59 | sean-k-mooney | but not for tox jobs | |
| 17:42:32 | sean-k-mooney | so if you see the tox py38 job fail that will be why | |
| 17:42:42 | sean-k-mooney | assuming your tests are otherwise correct :) | |
| 17:45:52 | sean-k-mooney | ricolin: the patch is deffently not correct but ill comment inline | |
| 17:46:27 | sean-k-mooney | ricolin: libvirt is never going to report a iommu model of auto or none | |
| 17:46:52 | sean-k-mooney | so you need to actully see what is reported form the domain caps api | |
| 17:47:35 | sean-k-mooney | by doing virsh domcapabilities --machine q35 --arch x86_64 | |
| 17:50:11 | sean-k-mooney | ricolin: but looking at that this is not somethign that is reported in that api | |
| 17:50:33 | sean-k-mooney | so instead of looking at the domaincap api you need to report the traits based on the libvirt version number | |
| 18:22:09 | melwitt | artom, sean-k-mooney: dunno if yall have seen this related preserve_on_delete bug from a few years ago https://bugs.launchpad.net/nova/+bug/1834463 | |
| 18:43:33 | opendevreview | Merged openstack/nova stable/ussuri: [stable-only] Make sdk broken job non voting until it is fixed https://review.opendev.org/c/openstack/nova/+/844309 | |
| 18:52:32 | ricolin | sean-k-mooney: so I need to check libvirt version before I put iommu in devices for fakelibvirt, right? | |
| 19:23:54 | artom | melwitt, hrmm, good find | |
| 19:43:24 | melwitt | artom: I looked through the code and saw that _heal_instance_info_cache preserves the existing value of preserve_on_delete. tried it out on devstack (created server with nova creating port, changed the value of preserve_on_delete to true in the database, saw _heal_instance_info_cache run a number of times, then detached the port) and it did not delete the port | |
| 19:49:24 | melwitt | I'm realizing the scenario in the above bug is different. they're saying they removed an interface by a manual database update, then nova added it back without (obviously) the original value of preserve_on_delete. I guess they are saying if they detach the port and then reattach it, they don't get the same value of preserve_on_delete. a bit different issue | |
| 19:56:45 | melwitt | although, they should get the same value bc if they reattach the port, nova won't consider it to be created by nova and thus should set preserve_on_delete = True | |
| 20:00:50 | sean-k-mooney | melwitt: they were updating the db | |
| 20:01:03 | sean-k-mooney | so really all bets are off at that point | |
| 20:01:24 | melwitt | just tried reattach and it indeed has preserve_on_delete = true. that means the bug report is very specifically the case where the interface gets removed from the info cache not via the API and then _heal_instance_info_cache runs. I don't know how that could happen during normal operation (no manual db update) | |
| 20:01:43 | sean-k-mooney | so there case does not make sense | |
| 20:01:47 | sean-k-mooney | well | |
| 20:02:04 | melwitt | yeah, I assumed they did the manual update to simplify a real world case but without any more data, I don't know how that case can happen | |
| 20:02:07 | sean-k-mooney | the booted with nova creating a nic | |
| 20:02:15 | sean-k-mooney | they somehow detached it without it gettting deleted | |
| 20:02:18 | sean-k-mooney | and then reattached it | |
| 20:02:41 | sean-k-mooney | so with the undocumented behavior when it got detach it shoudl have gotten delted | |
| 20:02:46 | sean-k-mooney | so there is not port to reattach | |
| 20:02:53 | melwitt | no, in their report they say they called server create with port_id passed in | |
| 20:03:12 | melwitt | so that means it begins with preserve_on_delete = true | |
| 20:03:15 | sean-k-mooney | oh then it shoudl have preserve on delete ture | |
| 20:03:23 | melwitt | yeah | |
| 20:03:37 | melwitt | no idea how what they say can happen "in real life" | |
| 20:04:01 | sean-k-mooney | so lest see | |
| 20:04:13 | sean-k-mooney | they are simulated the network info cache getting currpted | |
| 20:04:28 | sean-k-mooney | and then wating fo the heal taks to fix the info cache | |
| 20:04:39 | sean-k-mooney | and then nova things its created by it | |
| 20:05:00 | sean-k-mooney | i guess i can see that happeing if we lost the info of how the port was requested | |