| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-05-05 | |||
| 16:11:43 | dansmith | would it not be easier to just redefine all or most of our nodesets to be the nested label and then I could just run regular jobs depends-on that devstack change without needing to modify the nodeset everywhere? | |
| 16:12:30 | opendevreview | Artom Lifshitz proposed openstack/nova stable/wallaby: Reproduce bug 1995153 https://review.opendev.org/c/openstack/nova/+/882321 | |
| 16:12:31 | opendevreview | Artom Lifshitz proposed openstack/nova stable/wallaby: Save cell socket correctly when updating host NUMA topology https://review.opendev.org/c/openstack/nova/+/882322 | |
| 16:16:38 | clarkb | yes you could do that too | |
| 16:16:55 | dansmith | okay, lemme try that | |
| 16:20:04 | clarkb | it would probably be worth a note in the commit message that you shouldn't merge that change bceuase it will severely limit how many available nodes are available for running all devstack based jobs | |
| 16:20:10 | clarkb | I think it is a single cloud currently | |
| 16:20:15 | dansmith | yep, marked as DNM | |
| 16:20:43 | clarkb | dansmith: oh and you need to update devstack to not force emulation by default | |
| 16:20:56 | dansmith | yeah I saw that | |
| 16:20:59 | dansmith | I guess I better do that in the devstack patch too | |
| 16:21:01 | dansmith | I found it already | |
| 16:21:49 | clarkb | looks like devstack/.zuul.yaml grep for LIBVIRT_TYPE and set to kvm | |
| 16:21:53 | clarkb | but you found it | |
| 16:22:08 | clarkb | (it is set twice I think both need modification) | |
| 16:22:16 | dansmith | yu[p | |
| 16:26:20 | dansmith | clarkb: https://review.opendev.org/c/openstack/devstack/+/882457 | |
| 16:26:33 | dansmith | "does not match definition on master" ? | |
| 16:27:25 | clarkb | bah | |
| 16:27:41 | clarkb | I guess you will need to add definitions then with unique names | |
| 16:30:12 | dansmith | how would one ever change them then/ | |
| 16:30:22 | dansmith | define new and delete the old or something? | |
| 16:30:33 | clarkb | yes | |
| 16:31:17 | clarkb | there are alternatives as well. You can define anonymous nodesets in jobs directly (without names they just appl when that job runs) or in a central unbranched repo then there is a single copy of them (project-config could serve this purpose) | |
| 16:31:29 | clarkb | I believe they are defined directly in devstack to make it easier for third party ci systems to use though | |
| 16:31:37 | dansmith | okay so I should be able to just do this in c-t-p's .zuul then | |
| 16:31:50 | clarkb | yes that would work | |
| 16:31:52 | spatel | sean-k-mooney afternoon! | |
| 16:32:05 | clarkb | and then you can flip the LIBVIRT_TYPE var there too I think | |
| 16:32:42 | spatel | If you around then i have question related nova DB cleanup stuff, my machine oom out and some VMs stuck in nova DB and not sure how to clean them out. | |
| 16:48:52 | dansmith | clarkb: I'm messing up something with the nodeset definition: https://review.opendev.org/c/openstack/cinder-tempest-plugin/+/882458 | |
| 16:48:56 | dansmith | the error message is less helpful this time | |
| 16:50:57 | dansmith | oh wait | |
| 16:51:11 | dansmith | is it because name isn't indented? must be it | |
| 17:03:11 | dansmith | i believe it's running as expected now, thanks clarkb | |
| 17:34:54 | opendevreview | Balazs Gibizer proposed openstack/nova master: [doc]Clarify devname support in pci.device_spec https://review.opendev.org/c/openstack/nova/+/882464 | |
| 18:03:07 | clarkb | sorry I stepped out for a bit | |
| 18:07:31 | dansmith | clarkb: no problem I think I'm good now | |
| 18:07:46 | dansmith | clarkb: semi-related, am I just stupid or is it impossible to get two conditions in the opensearch query? | |
| 18:08:01 | dansmith | any two-condition search I do always returns no results | |
| 18:08:23 | dansmith | and I get a syntax error | |
| 18:08:56 | dansmith | oh I guess I need an "and" operator | |
| 18:11:07 | clarkb | I don't know I'm not really invovled in it | |
| 18:12:05 | dansmith | it seems like I used to be able to stumble myself into a useful query and since the upgrade I can only get really basic stuff to work | |
| 18:44:01 | dansmith | clarkb: so, I guess something is still wrong because lib/nova switches my requested kvm to qemu because /dev/kvm is not accessible | |
| 18:44:28 | dansmith | https://github.com/openstack/devstack/blob/master/lib/nova#L269 | |
| 18:44:59 | dansmith | oh, because that job didn't select the nested-jammy label, hrm | |
| 19:36:30 | spatel | Any idea how to clean up orphan VMs entries in nova DB? | |
| 19:36:54 | spatel | I have used virsh destroy command to delete vms and now DB has entries for them but VM doesn't exist | |
| 19:40:52 | dansmith | spatel: virsh destroy does nothing for nova vms, nova will just try to recreate them | |
| 19:41:32 | spatel | Hmm, I did delete them in openstack also using nova delete command | |
| 19:42:08 | dansmith | that's the only way, but they remain in the database until you archive (as you noted in your mailing list post) | |
| 19:42:20 | dansmith | archive will only remove them if they're marked as deleted | |
| 19:43:05 | spatel | They are doesn't exist https://paste.opendev.org/show/b0Caj0S65vx4hBFhAyML/ | |
| 19:43:32 | spatel | openstack hypervisor stat showing 91 vms running but openstack servers list showing only single VM | |
| 19:43:51 | spatel | Definitely nova DB is out of sync | |
| 19:44:30 | spatel | I look into nova/instances DB table and there is only single entry | |
| 19:45:14 | dansmith | then what's the problem? | |
| 19:45:32 | dansmith | just the running_vms count? | |
| 19:45:36 | spatel | Yes... | |
| 19:45:58 | spatel | How do i may everything in sync ? | |
| 19:46:30 | spatel | Just curious from where openstack hypervisor command finding 91 vms? | |
| 19:47:05 | dansmith | you need to look at compute_nodes.running_vms to see which one is still reporting instances | |
| 19:47:31 | spatel | let me take a look at that table | |
| 19:51:06 | spatel | I am not able to find that tables in DB | |
| 19:51:25 | spatel | it should be inside nova/instance correct? | |
| 19:52:07 | dansmith | instance (assuming you meant instances) is a table, compute_nodes is a table, running_vms is a column in the compute_nodes table | |
| 19:53:23 | spatel | found it | |
| 19:56:00 | spatel | Yes, I can see them there that on node1 - https://paste.opendev.org/show/bnhl6ZeHXIA3Tjk1ucV6/ | |
| 19:56:51 | spatel | do you think just update those number in table is enough? | |
| 19:56:55 | dansmith | that's three nodes | |
| 19:56:57 | dansmith | no | |
| 19:57:05 | dansmith | I mean, that will make the number change, but it's not the right fix | |
| 19:57:17 | dansmith | you need to select the hostname along with the count to know which is which | |
| 19:57:27 | dansmith | nova-compute should be updating those numbers | |
| 19:58:19 | dansmith | select host,hypervisor_hostname,running_vms from compute_nodes; | |
| 19:58:48 | spatel | https://paste.opendev.org/show/bwohGVNFtteg3J4x5qUY/ | |
| 19:59:10 | spatel | at present on ctrl node there are no VM running.. | |
| 19:59:22 | spatel | at present on ctrl1 and ctrl3 node there are no VM running.. | |
| 19:59:37 | spatel | That entry should be zero technically | |
| 19:59:50 | dansmith | is nova-compute running on each of those three nodes? | |
| 19:59:59 | dansmith | because it should be updating that number every few minutes | |
| 20:00:03 | spatel | yes its running | |
| 20:00:30 | spatel | all services showing fine.. I have restarted them | |
| 20:00:32 | spatel | no nasty logs or errors anywhere | |
| 20:02:50 | dansmith | they should all be iterating over their instances regularly and updating those numbers | |
| 20:05:12 | dansmith | perhaps it's not doing that if there are no instances (although you said there was one, so at least that one should be correct) | |
| 20:06:58 | spatel | out of 3 nodes only node2 has 1 VM running and rest are empty | |
| 20:07:41 | dansmith | yeah, so that node should show 1 in the database and doesn't, which to me means something is wrong (unless it hasn't run update_available_resource yet) | |
| 20:07:43 | spatel | Thinking to reboot all 3 nodes to start with fresh troubleshooting | |
| 20:08:30 | spatel | who will run update_available_resource task? compute nodes correct? | |
| 20:08:53 | dansmith | nova-compute does it | |
| 20:09:11 | spatel | may be rabbitMQ is in zombie state... I have checked cluster_status and its showing all good but who knows.. | |
| 20:09:47 | dansmith | there should be errors in nova-compute if so, but hard to say after something like an oom | |
| 20:11:26 | spatel | Let me check.. | |
| 20:12:58 | spatel | I found this lines in nova-compute - AMQP server on 192.168.1.11:5672 is unreachable: timed out. Trying again in 0 seconds.: socket.timeout: timed out | |
| 20:13:11 | spatel | Look like issue is related to rabbit.. hmm | |
| 20:13:22 | spatel | but cluster status is green | |
| 20:13:56 | spatel | Let me destroy rabbit and rebuild it to see if it come clean | |
| 20:27:15 | spatel | dansmith look like it was rabbit issue, after re-building rabbit I can see correct count on hypervisor stats :) | |