Earlier  
Posted Nick Remark
#openstack-nova - 2023-05-05
16:10:13 clarkb It just needs to follow the same basic format of the devstack nodesets because devstack has expectations about groups. You can change the nodeset name and node labels though
16:10:57 dansmith ah, okay
16:11:43 dansmith would it not be easier to just redefine all or most of our nodesets to be the nested label and then I could just run regular jobs depends-on that devstack change without needing to modify the nodeset everywhere?
16:12:30 opendevreview Artom Lifshitz proposed openstack/nova stable/wallaby: Reproduce bug 1995153 https://review.opendev.org/c/openstack/nova/+/882321
16:12:31 opendevreview Artom Lifshitz proposed openstack/nova stable/wallaby: Save cell socket correctly when updating host NUMA topology https://review.opendev.org/c/openstack/nova/+/882322
16:16:38 clarkb yes you could do that too
16:16:55 dansmith okay, lemme try that
16:20:04 clarkb it would probably be worth a note in the commit message that you shouldn't merge that change bceuase it will severely limit how many available nodes are available for running all devstack based jobs
16:20:10 clarkb I think it is a single cloud currently
16:20:15 dansmith yep, marked as DNM
16:20:43 clarkb dansmith: oh and you need to update devstack to not force emulation by default
16:20:56 dansmith yeah I saw that
16:20:59 dansmith I guess I better do that in the devstack patch too
16:21:01 dansmith I found it already
16:21:49 clarkb looks like devstack/.zuul.yaml grep for LIBVIRT_TYPE and set to kvm
16:21:53 clarkb but you found it
16:22:08 clarkb (it is set twice I think both need modification)
16:22:16 dansmith yu[p
16:26:20 dansmith clarkb: https://review.opendev.org/c/openstack/devstack/+/882457
16:26:33 dansmith "does not match definition on master" ?
16:27:25 clarkb bah
16:27:41 clarkb I guess you will need to add definitions then with unique names
16:30:12 dansmith how would one ever change them then/
16:30:22 dansmith define new and delete the old or something?
16:30:33 clarkb yes
16:31:17 clarkb there are alternatives as well. You can define anonymous nodesets in jobs directly (without names they just appl when that job runs) or in a central unbranched repo then there is a single copy of them (project-config could serve this purpose)
16:31:29 clarkb I believe they are defined directly in devstack to make it easier for third party ci systems to use though
16:31:37 dansmith okay so I should be able to just do this in c-t-p's .zuul then
16:31:50 clarkb yes that would work
16:31:52 spatel sean-k-mooney afternoon!
16:32:05 clarkb and then you can flip the LIBVIRT_TYPE var there too I think
16:32:42 spatel If you around then i have question related nova DB cleanup stuff, my machine oom out and some VMs stuck in nova DB and not sure how to clean them out.
16:48:52 dansmith clarkb: I'm messing up something with the nodeset definition: https://review.opendev.org/c/openstack/cinder-tempest-plugin/+/882458
16:48:56 dansmith the error message is less helpful this time
16:50:57 dansmith oh wait
16:51:11 dansmith is it because name isn't indented? must be it
17:03:11 dansmith i believe it's running as expected now, thanks clarkb
17:34:54 opendevreview Balazs Gibizer proposed openstack/nova master: [doc]Clarify devname support in pci.device_spec https://review.opendev.org/c/openstack/nova/+/882464
18:03:07 clarkb sorry I stepped out for a bit
18:07:31 dansmith clarkb: no problem I think I'm good now
18:07:46 dansmith clarkb: semi-related, am I just stupid or is it impossible to get two conditions in the opensearch query?
18:08:01 dansmith any two-condition search I do always returns no results
18:08:23 dansmith and I get a syntax error
18:08:56 dansmith oh I guess I need an "and" operator
18:11:07 clarkb I don't know I'm not really invovled in it
18:12:05 dansmith it seems like I used to be able to stumble myself into a useful query and since the upgrade I can only get really basic stuff to work
18:44:01 dansmith clarkb: so, I guess something is still wrong because lib/nova switches my requested kvm to qemu because /dev/kvm is not accessible
18:44:28 dansmith https://github.com/openstack/devstack/blob/master/lib/nova#L269
18:44:59 dansmith oh, because that job didn't select the nested-jammy label, hrm
19:36:30 spatel Any idea how to clean up orphan VMs entries in nova DB?
19:36:54 spatel I have used virsh destroy command to delete vms and now DB has entries for them but VM doesn't exist
19:40:52 dansmith spatel: virsh destroy does nothing for nova vms, nova will just try to recreate them
19:41:32 spatel Hmm, I did delete them in openstack also using nova delete command
19:42:08 dansmith that's the only way, but they remain in the database until you archive (as you noted in your mailing list post)
19:42:20 dansmith archive will only remove them if they're marked as deleted
19:43:05 spatel They are doesn't exist https://paste.opendev.org/show/b0Caj0S65vx4hBFhAyML/
19:43:32 spatel openstack hypervisor stat showing 91 vms running but openstack servers list showing only single VM
19:43:51 spatel Definitely nova DB is out of sync
19:44:30 spatel I look into nova/instances DB table and there is only single entry
19:45:14 dansmith then what's the problem?
19:45:32 dansmith just the running_vms count?
19:45:36 spatel Yes...
19:45:58 spatel How do i may everything in sync ?
19:46:30 spatel Just curious from where openstack hypervisor command finding 91 vms?
19:47:05 dansmith you need to look at compute_nodes.running_vms to see which one is still reporting instances
19:47:31 spatel let me take a look at that table
19:51:06 spatel I am not able to find that tables in DB
19:51:25 spatel it should be inside nova/instance correct?
19:52:07 dansmith instance (assuming you meant instances) is a table, compute_nodes is a table, running_vms is a column in the compute_nodes table
19:53:23 spatel found it
19:56:00 spatel Yes, I can see them there that on node1 - https://paste.opendev.org/show/bnhl6ZeHXIA3Tjk1ucV6/
19:56:51 spatel do you think just update those number in table is enough?
19:56:55 dansmith that's three nodes
19:56:57 dansmith no
19:57:05 dansmith I mean, that will make the number change, but it's not the right fix
19:57:17 dansmith you need to select the hostname along with the count to know which is which
19:57:27 dansmith nova-compute should be updating those numbers
19:58:19 dansmith select host,hypervisor_hostname,running_vms from compute_nodes;
19:58:48 spatel https://paste.opendev.org/show/bwohGVNFtteg3J4x5qUY/
19:59:10 spatel at present on ctrl node there are no VM running..
19:59:22 spatel at present on ctrl1 and ctrl3 node there are no VM running..
19:59:37 spatel That entry should be zero technically
19:59:50 dansmith is nova-compute running on each of those three nodes?
19:59:59 dansmith because it should be updating that number every few minutes
20:00:03 spatel yes its running
20:00:30 spatel all services showing fine.. I have restarted them
20:00:32 spatel no nasty logs or errors anywhere
20:02:50 dansmith they should all be iterating over their instances regularly and updating those numbers
20:05:12 dansmith perhaps it's not doing that if there are no instances (although you said there was one, so at least that one should be correct)
20:06:58 spatel out of 3 nodes only node2 has 1 VM running and rest are empty
20:07:41 dansmith yeah, so that node should show 1 in the database and doesn't, which to me means something is wrong (unless it hasn't run update_available_resource yet)
20:07:43 spatel Thinking to reboot all 3 nodes to start with fresh troubleshooting
20:08:30 spatel who will run update_available_resource task? compute nodes correct?
20:08:53 dansmith nova-compute does it
20:09:11 spatel may be rabbitMQ is in zombie state... I have checked cluster_status and its showing all good but who knows..
20:09:47 dansmith there should be errors in nova-compute if so, but hard to say after something like an oom
20:11:26 spatel Let me check..
20:12:58 spatel I found this lines in nova-compute - AMQP server on 192.168.1.11:5672 is unreachable: timed out. Trying again in 0 seconds.: socket.timeout: timed out
20:13:11 spatel Look like issue is related to rabbit.. hmm
20:13:22 spatel but cluster status is green

Earlier   Later