Earlier  
Posted Nick Remark
#openstack-nova - 2023-05-05
16:16:38 clarkb yes you could do that too
16:16:55 dansmith okay, lemme try that
16:20:04 clarkb it would probably be worth a note in the commit message that you shouldn't merge that change bceuase it will severely limit how many available nodes are available for running all devstack based jobs
16:20:10 clarkb I think it is a single cloud currently
16:20:15 dansmith yep, marked as DNM
16:20:43 clarkb dansmith: oh and you need to update devstack to not force emulation by default
16:20:56 dansmith yeah I saw that
16:20:59 dansmith I guess I better do that in the devstack patch too
16:21:01 dansmith I found it already
16:21:49 clarkb looks like devstack/.zuul.yaml grep for LIBVIRT_TYPE and set to kvm
16:21:53 clarkb but you found it
16:22:08 clarkb (it is set twice I think both need modification)
16:22:16 dansmith yu[p
16:26:20 dansmith clarkb: https://review.opendev.org/c/openstack/devstack/+/882457
16:26:33 dansmith "does not match definition on master" ?
16:27:25 clarkb bah
16:27:41 clarkb I guess you will need to add definitions then with unique names
16:30:12 dansmith how would one ever change them then/
16:30:22 dansmith define new and delete the old or something?
16:30:33 clarkb yes
16:31:17 clarkb there are alternatives as well. You can define anonymous nodesets in jobs directly (without names they just appl when that job runs) or in a central unbranched repo then there is a single copy of them (project-config could serve this purpose)
16:31:29 clarkb I believe they are defined directly in devstack to make it easier for third party ci systems to use though
16:31:37 dansmith okay so I should be able to just do this in c-t-p's .zuul then
16:31:50 clarkb yes that would work
16:31:52 spatel sean-k-mooney afternoon!
16:32:05 clarkb and then you can flip the LIBVIRT_TYPE var there too I think
16:32:42 spatel If you around then i have question related nova DB cleanup stuff, my machine oom out and some VMs stuck in nova DB and not sure how to clean them out.
16:48:52 dansmith clarkb: I'm messing up something with the nodeset definition: https://review.opendev.org/c/openstack/cinder-tempest-plugin/+/882458
16:48:56 dansmith the error message is less helpful this time
16:50:57 dansmith oh wait
16:51:11 dansmith is it because name isn't indented? must be it
17:03:11 dansmith i believe it's running as expected now, thanks clarkb
17:34:54 opendevreview Balazs Gibizer proposed openstack/nova master: [doc]Clarify devname support in pci.device_spec https://review.opendev.org/c/openstack/nova/+/882464
18:03:07 clarkb sorry I stepped out for a bit
18:07:31 dansmith clarkb: no problem I think I'm good now
18:07:46 dansmith clarkb: semi-related, am I just stupid or is it impossible to get two conditions in the opensearch query?
18:08:01 dansmith any two-condition search I do always returns no results
18:08:23 dansmith and I get a syntax error
18:08:56 dansmith oh I guess I need an "and" operator
18:11:07 clarkb I don't know I'm not really invovled in it
18:12:05 dansmith it seems like I used to be able to stumble myself into a useful query and since the upgrade I can only get really basic stuff to work
18:44:01 dansmith clarkb: so, I guess something is still wrong because lib/nova switches my requested kvm to qemu because /dev/kvm is not accessible
18:44:28 dansmith https://github.com/openstack/devstack/blob/master/lib/nova#L269
18:44:59 dansmith oh, because that job didn't select the nested-jammy label, hrm
19:36:30 spatel Any idea how to clean up orphan VMs entries in nova DB?
19:36:54 spatel I have used virsh destroy command to delete vms and now DB has entries for them but VM doesn't exist
19:40:52 dansmith spatel: virsh destroy does nothing for nova vms, nova will just try to recreate them
19:41:32 spatel Hmm, I did delete them in openstack also using nova delete command
19:42:08 dansmith that's the only way, but they remain in the database until you archive (as you noted in your mailing list post)
19:42:20 dansmith archive will only remove them if they're marked as deleted
19:43:05 spatel They are doesn't exist https://paste.opendev.org/show/b0Caj0S65vx4hBFhAyML/
19:43:32 spatel openstack hypervisor stat showing 91 vms running but openstack servers list showing only single VM
19:43:51 spatel Definitely nova DB is out of sync
19:44:30 spatel I look into nova/instances DB table and there is only single entry
19:45:14 dansmith then what's the problem?
19:45:32 dansmith just the running_vms count?
19:45:36 spatel Yes...
19:45:58 spatel How do i may everything in sync ?
19:46:30 spatel Just curious from where openstack hypervisor command finding 91 vms?
19:47:05 dansmith you need to look at compute_nodes.running_vms to see which one is still reporting instances
19:47:31 spatel let me take a look at that table
19:51:06 spatel I am not able to find that tables in DB
19:51:25 spatel it should be inside nova/instance correct?
19:52:07 dansmith instance (assuming you meant instances) is a table, compute_nodes is a table, running_vms is a column in the compute_nodes table
19:53:23 spatel found it
19:56:00 spatel Yes, I can see them there that on node1 - https://paste.opendev.org/show/bnhl6ZeHXIA3Tjk1ucV6/
19:56:51 spatel do you think just update those number in table is enough?
19:56:55 dansmith that's three nodes
19:56:57 dansmith no
19:57:05 dansmith I mean, that will make the number change, but it's not the right fix
19:57:17 dansmith you need to select the hostname along with the count to know which is which
19:57:27 dansmith nova-compute should be updating those numbers
19:58:19 dansmith select host,hypervisor_hostname,running_vms from compute_nodes;
19:58:48 spatel https://paste.opendev.org/show/bwohGVNFtteg3J4x5qUY/
19:59:10 spatel at present on ctrl node there are no VM running..
19:59:22 spatel at present on ctrl1 and ctrl3 node there are no VM running..
19:59:37 spatel That entry should be zero technically
19:59:50 dansmith is nova-compute running on each of those three nodes?
19:59:59 dansmith because it should be updating that number every few minutes
20:00:03 spatel yes its running
20:00:30 spatel all services showing fine.. I have restarted them
20:00:32 spatel no nasty logs or errors anywhere
20:02:50 dansmith they should all be iterating over their instances regularly and updating those numbers
20:05:12 dansmith perhaps it's not doing that if there are no instances (although you said there was one, so at least that one should be correct)
20:06:58 spatel out of 3 nodes only node2 has 1 VM running and rest are empty
20:07:41 dansmith yeah, so that node should show 1 in the database and doesn't, which to me means something is wrong (unless it hasn't run update_available_resource yet)
20:07:43 spatel Thinking to reboot all 3 nodes to start with fresh troubleshooting
20:08:30 spatel who will run update_available_resource task? compute nodes correct?
20:08:53 dansmith nova-compute does it
20:09:11 spatel may be rabbitMQ is in zombie state... I have checked cluster_status and its showing all good but who knows..
20:09:47 dansmith there should be errors in nova-compute if so, but hard to say after something like an oom
20:11:26 spatel Let me check..
20:12:58 spatel I found this lines in nova-compute - AMQP server on 192.168.1.11:5672 is unreachable: timed out. Trying again in 0 seconds.: socket.timeout: timed out
20:13:11 spatel Look like issue is related to rabbit.. hmm
20:13:22 spatel but cluster status is green
20:13:56 spatel Let me destroy rabbit and rebuild it to see if it come clean
20:27:15 spatel dansmith look like it was rabbit issue, after re-building rabbit I can see correct count on hypervisor stats :)
20:28:00 dansmith spatel: good, that's why I was recommending you not just fix it manually because it's an indication of something else
20:28:48 spatel Thank you for staying with me :)
20:28:56 spatel I was about to blowup DB.. haha

Earlier   Later