Earlier  
Posted Nick Remark
#openstack-nova - 2022-06-02
15:32:17 kashyap slaweq: A quick question: the instance simply crashes when launching it?
15:32:51 slaweq kashyap I think it crashed during snapshoting
15:32:55 slaweq it was spawned properly
15:33:00 opendevreview Alexey Stupnikov proposed openstack/nova master: Optimize _local_delete calls by compute unit tests https://review.opendev.org/c/openstack/nova/+/844285
15:33:51 dansmith kashyap: ah, running it on a specific pid would be much harder
15:34:26 dansmith running it like you describe is something we could hack in as a post job
15:34:32 kashyap dansmith: Nah, we can disregad the per-PID thing
15:34:47 kashyap Yeah, binary is easier indeed
15:35:23 dansmith okay after call(s) I can help hack that in if we need, but if it's pretty repeatable it might be easier to just try to repro locally
15:36:01 kashyap dansmith: Yeah, that's the next thing I'm looking. It looks like it's not any Ceph-based, and just plain local storage, IIRC
15:36:08 dansmith cool
15:37:49 kashyap slaweq: I'm just trying to find the precise test trigger. From looking at the 'n-cpu' log, snapshots seem to happen just fine. I'll look more after I'm done w/ this call
15:38:27 slaweq kashyap but IIUC nova logs, instance is gone during snapshoting process
15:38:37 slaweq please take Your time, it's not urgent for us for sure
15:38:52 dansmith kashyap: it's test_create_backup
15:38:55 dansmith https://storage.bhs.cloud.ovh.net/v1/AUTH_dcaab5e32b234d56b626f72581e3644c/zuul_opendev_logs_4a7/periodic/opendev.org/openstack/neutron/master/neutron-ovn-tempest-ovs-master-fedora/4a7f284/testr_results.html
15:39:25 kashyap (Yeah, just found it; thx)
15:43:08 kashyap dansmith: Thanks; so the above command I noted above "coredumpctl dump" will dump the most recent core dump. I guess we have to redirect it to a file
15:43:11 opendevreview Merged openstack/osc-placement master: Add Python3 zed unit tests https://review.opendev.org/c/openstack/osc-placement/+/835369
15:43:39 kashyap dansmith: Err, ignore the above comment; the "-o" is the file.
15:43:58 dansmith kashyap: ack, are you going to try to repro locally first?
15:44:33 kashyap dansmith: On F35, just run the Tempest test, or construct a manual libvirt-based repro?
15:44:37 dansmith if it doesn't repro locally that might be interesting to know as well, like whether it's related to having run a lot of tests first, or that one thing always fails in isolation
15:44:55 dansmith kashyap: I just meant devstack, run that one tempest test
15:45:33 kashyap Ah; nod. I can't today; but I can give it a go tomorrow.
15:45:37 dansmith when I'm done here I can work on adding that as a post task, it just might take a bunch of iterations to get it right (based on experience)
15:45:45 dansmith okay, I'll give it a shot at least when I'm done here
15:47:00 kashyap dansmith: When you say "adding that as a post task" -- I take it you mean adding the above "coredumpctl ... dump", yeah?
15:47:08 dansmith yep
15:49:56 kashyap I have a deja vu about this test_create_backup test, reading its code
15:56:14 kashyap slaweq: When you get a minute, can you please file an upstream LP bug to track this? So we can keep all the investigation in one place?
16:01:12 dansmith kashyap: against what, nova?
16:01:26 kashyap Yeah, I'd say so
16:01:33 kashyap dansmith: I just looked at the compressed libvirtd.log
16:01:46 kashyap And I see a familiar libvirt error:
16:01:46 kashyap 2022-06-01 03:35:33.685+0000: 87576: error : qemuMonitorJSONCheckErrorFull:412 : internal error: unable to execute QEMU command 'blockdev-del': Failed to find node with node-name='libvirt-5-storage'
16:02:30 kashyap dansmith: In the past we found the same error earlier this year, and I recall working w/ libvirt folks to get a fix. But that was in a different context: https://listman.redhat.com/archives/libvir-list/2022-February/msg00790.html
16:03:11 kashyap dansmith: slaweq: The root TripleO (actually should've filed for Nova) bug was this where did the analysis: https://bugs.launchpad.net/tripleo/+bug/1959014
16:03:26 kashyap If you open that last link, scroll from bottom for more signal
16:54:24 dansmith kashyap: yeah I remember that one.. so to be clear, you expect this is a different issue right?
17:09:24 dansmith kashyap: this is running, we'll see: https://review.opendev.org/c/openstack/devstack/+/844503
17:10:56 ricolin bauzas: I think https://review.opendev.org/c/openstack/nova/+/830646 is ready for review now, could you kindly remove the -2
17:11:54 bauzas ricolin: sure, lemme look
17:12:13 ricolin bauzas: thanks:)
17:12:33 bauzas ricolin: oh, yeah you created the bp and the spec, ta
17:12:57 ricolin bauzas: yeah, the spec merged:)
17:13:24 ricolin and I updated the implement patch accordingly
17:13:32 ricolin I think:)
17:20:05 opendevreview Rico Lin proposed openstack/nova master: Add traits for viommu model https://review.opendev.org/c/openstack/nova/+/844507
17:21:17 opendevreview Artom Lifshitz proposed openstack/nova stable/ussuri: fake: Ensure need_legacy_block_device_info returns False https://review.opendev.org/c/openstack/nova/+/843950
17:21:18 opendevreview Artom Lifshitz proposed openstack/nova stable/ussuri: Add a regression test for bug 1939545 https://review.opendev.org/c/openstack/nova/+/843951
17:21:19 opendevreview Artom Lifshitz proposed openstack/nova stable/ussuri: compute: Ensure updates to bdms during pre_live_migration are saved https://review.opendev.org/c/openstack/nova/+/843952
17:24:03 ricolin sean-k-mooney: this should be the last piece for libvirt-viommu-device implementation, but as I'm not familiar with traits, can you take a review on it and let me know if I do it right/wrong
17:24:05 ricolin https://review.opendev.org/c/openstack/nova/+/844507
17:41:15 sean-k-mooney sure
17:41:30 sean-k-mooney just so you are aware the unit test will fail untill the trait is merged and released
17:41:50 sean-k-mooney but the tempest test shoudl be able to pass becauses depens on works for devstack jobs
17:41:59 sean-k-mooney but not for tox jobs
17:42:32 sean-k-mooney so if you see the tox py38 job fail that will be why
17:42:42 sean-k-mooney assuming your tests are otherwise correct :)
17:45:52 sean-k-mooney ricolin: the patch is deffently not correct but ill comment inline
17:46:27 sean-k-mooney ricolin: libvirt is never going to report a iommu model of auto or none
17:46:52 sean-k-mooney so you need to actully see what is reported form the domain caps api
17:47:35 sean-k-mooney by doing virsh domcapabilities --machine q35 --arch x86_64
17:50:11 sean-k-mooney ricolin: but looking at that this is not somethign that is reported in that api
17:50:33 sean-k-mooney so instead of looking at the domaincap api you need to report the traits based on the libvirt version number
18:22:09 melwitt artom, sean-k-mooney: dunno if yall have seen this related preserve_on_delete bug from a few years ago https://bugs.launchpad.net/nova/+bug/1834463
18:43:33 opendevreview Merged openstack/nova stable/ussuri: [stable-only] Make sdk broken job non voting until it is fixed https://review.opendev.org/c/openstack/nova/+/844309
18:52:32 ricolin sean-k-mooney: so I need to check libvirt version before I put iommu in devices for fakelibvirt, right?
19:23:54 artom melwitt, hrmm, good find
19:43:24 melwitt artom: I looked through the code and saw that _heal_instance_info_cache preserves the existing value of preserve_on_delete. tried it out on devstack (created server with nova creating port, changed the value of preserve_on_delete to true in the database, saw _heal_instance_info_cache run a number of times, then detached the port) and it did not delete the port
19:49:24 melwitt I'm realizing the scenario in the above bug is different. they're saying they removed an interface by a manual database update, then nova added it back without (obviously) the original value of preserve_on_delete. I guess they are saying if they detach the port and then reattach it, they don't get the same value of preserve_on_delete. a bit different issue
19:56:45 melwitt although, they should get the same value bc if they reattach the port, nova won't consider it to be created by nova and thus should set preserve_on_delete = True
20:00:50 sean-k-mooney melwitt: they were updating the db
20:01:03 sean-k-mooney so really all bets are off at that point
20:01:24 melwitt just tried reattach and it indeed has preserve_on_delete = true. that means the bug report is very specifically the case where the interface gets removed from the info cache not via the API and then _heal_instance_info_cache runs. I don't know how that could happen during normal operation (no manual db update)
20:01:43 sean-k-mooney so there case does not make sense
20:01:47 sean-k-mooney well
20:02:04 melwitt yeah, I assumed they did the manual update to simplify a real world case but without any more data, I don't know how that case can happen
20:02:07 sean-k-mooney the booted with nova creating a nic
20:02:15 sean-k-mooney they somehow detached it without it gettting deleted
20:02:18 sean-k-mooney and then reattached it
20:02:41 sean-k-mooney so with the undocumented behavior when it got detach it shoudl have gotten delted
20:02:46 sean-k-mooney so there is not port to reattach
20:02:53 melwitt no, in their report they say they called server create with port_id passed in
20:03:12 melwitt so that means it begins with preserve_on_delete = true
20:03:15 sean-k-mooney oh then it shoudl have preserve on delete ture
20:03:23 melwitt yeah
20:03:37 melwitt no idea how what they say can happen "in real life"
20:04:01 sean-k-mooney so lest see
20:04:13 sean-k-mooney they are simulated the network info cache getting currpted
20:04:28 sean-k-mooney and then wating fo the heal taks to fix the info cache
20:04:39 sean-k-mooney and then nova things its created by it
20:05:00 sean-k-mooney i guess i can see that happeing if we lost the info of how the port was requested
20:05:28 sean-k-mooney so that is implying we sotre that in the info cache only
20:05:42 melwitt yeah, nova just sets the flag to true if it created the port, at port creation time. after that it's cache only
20:06:07 sean-k-mooney well thats broken
20:06:11 sean-k-mooney i guess we do that for attach
20:06:13 sean-k-mooney too
20:06:26 sean-k-mooney e.g. if we do attach network instead of attch port

Earlier   Later