| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-05-05 | |||
| 16:31:37 | dansmith | okay so I should be able to just do this in c-t-p's .zuul then | |
| 16:31:50 | clarkb | yes that would work | |
| 16:31:52 | spatel | sean-k-mooney afternoon! | |
| 16:32:05 | clarkb | and then you can flip the LIBVIRT_TYPE var there too I think | |
| 16:32:42 | spatel | If you around then i have question related nova DB cleanup stuff, my machine oom out and some VMs stuck in nova DB and not sure how to clean them out. | |
| 16:48:52 | dansmith | clarkb: I'm messing up something with the nodeset definition: https://review.opendev.org/c/openstack/cinder-tempest-plugin/+/882458 | |
| 16:48:56 | dansmith | the error message is less helpful this time | |
| 16:50:57 | dansmith | oh wait | |
| 16:51:11 | dansmith | is it because name isn't indented? must be it | |
| 17:03:11 | dansmith | i believe it's running as expected now, thanks clarkb | |
| 17:34:54 | opendevreview | Balazs Gibizer proposed openstack/nova master: [doc]Clarify devname support in pci.device_spec https://review.opendev.org/c/openstack/nova/+/882464 | |
| 18:03:07 | clarkb | sorry I stepped out for a bit | |
| 18:07:31 | dansmith | clarkb: no problem I think I'm good now | |
| 18:07:46 | dansmith | clarkb: semi-related, am I just stupid or is it impossible to get two conditions in the opensearch query? | |
| 18:08:01 | dansmith | any two-condition search I do always returns no results | |
| 18:08:23 | dansmith | and I get a syntax error | |
| 18:08:56 | dansmith | oh I guess I need an "and" operator | |
| 18:11:07 | clarkb | I don't know I'm not really invovled in it | |
| 18:12:05 | dansmith | it seems like I used to be able to stumble myself into a useful query and since the upgrade I can only get really basic stuff to work | |
| 18:44:01 | dansmith | clarkb: so, I guess something is still wrong because lib/nova switches my requested kvm to qemu because /dev/kvm is not accessible | |
| 18:44:28 | dansmith | https://github.com/openstack/devstack/blob/master/lib/nova#L269 | |
| 18:44:59 | dansmith | oh, because that job didn't select the nested-jammy label, hrm | |
| 19:36:30 | spatel | Any idea how to clean up orphan VMs entries in nova DB? | |
| 19:36:54 | spatel | I have used virsh destroy command to delete vms and now DB has entries for them but VM doesn't exist | |
| 19:40:52 | dansmith | spatel: virsh destroy does nothing for nova vms, nova will just try to recreate them | |
| 19:41:32 | spatel | Hmm, I did delete them in openstack also using nova delete command | |
| 19:42:08 | dansmith | that's the only way, but they remain in the database until you archive (as you noted in your mailing list post) | |
| 19:42:20 | dansmith | archive will only remove them if they're marked as deleted | |
| 19:43:05 | spatel | They are doesn't exist https://paste.opendev.org/show/b0Caj0S65vx4hBFhAyML/ | |
| 19:43:32 | spatel | openstack hypervisor stat showing 91 vms running but openstack servers list showing only single VM | |
| 19:43:51 | spatel | Definitely nova DB is out of sync | |
| 19:44:30 | spatel | I look into nova/instances DB table and there is only single entry | |
| 19:45:14 | dansmith | then what's the problem? | |
| 19:45:32 | dansmith | just the running_vms count? | |
| 19:45:36 | spatel | Yes... | |
| 19:45:58 | spatel | How do i may everything in sync ? | |
| 19:46:30 | spatel | Just curious from where openstack hypervisor command finding 91 vms? | |
| 19:47:05 | dansmith | you need to look at compute_nodes.running_vms to see which one is still reporting instances | |
| 19:47:31 | spatel | let me take a look at that table | |
| 19:51:06 | spatel | I am not able to find that tables in DB | |
| 19:51:25 | spatel | it should be inside nova/instance correct? | |
| 19:52:07 | dansmith | instance (assuming you meant instances) is a table, compute_nodes is a table, running_vms is a column in the compute_nodes table | |
| 19:53:23 | spatel | found it | |
| 19:56:00 | spatel | Yes, I can see them there that on node1 - https://paste.opendev.org/show/bnhl6ZeHXIA3Tjk1ucV6/ | |
| 19:56:51 | spatel | do you think just update those number in table is enough? | |
| 19:56:55 | dansmith | that's three nodes | |
| 19:56:57 | dansmith | no | |
| 19:57:05 | dansmith | I mean, that will make the number change, but it's not the right fix | |
| 19:57:17 | dansmith | you need to select the hostname along with the count to know which is which | |
| 19:57:27 | dansmith | nova-compute should be updating those numbers | |
| 19:58:19 | dansmith | select host,hypervisor_hostname,running_vms from compute_nodes; | |
| 19:58:48 | spatel | https://paste.opendev.org/show/bwohGVNFtteg3J4x5qUY/ | |
| 19:59:10 | spatel | at present on ctrl node there are no VM running.. | |
| 19:59:22 | spatel | at present on ctrl1 and ctrl3 node there are no VM running.. | |
| 19:59:37 | spatel | That entry should be zero technically | |
| 19:59:50 | dansmith | is nova-compute running on each of those three nodes? | |
| 19:59:59 | dansmith | because it should be updating that number every few minutes | |
| 20:00:03 | spatel | yes its running | |
| 20:00:30 | spatel | all services showing fine.. I have restarted them | |
| 20:00:32 | spatel | no nasty logs or errors anywhere | |
| 20:02:50 | dansmith | they should all be iterating over their instances regularly and updating those numbers | |
| 20:05:12 | dansmith | perhaps it's not doing that if there are no instances (although you said there was one, so at least that one should be correct) | |
| 20:06:58 | spatel | out of 3 nodes only node2 has 1 VM running and rest are empty | |
| 20:07:41 | dansmith | yeah, so that node should show 1 in the database and doesn't, which to me means something is wrong (unless it hasn't run update_available_resource yet) | |
| 20:07:43 | spatel | Thinking to reboot all 3 nodes to start with fresh troubleshooting | |
| 20:08:30 | spatel | who will run update_available_resource task? compute nodes correct? | |
| 20:08:53 | dansmith | nova-compute does it | |
| 20:09:11 | spatel | may be rabbitMQ is in zombie state... I have checked cluster_status and its showing all good but who knows.. | |
| 20:09:47 | dansmith | there should be errors in nova-compute if so, but hard to say after something like an oom | |
| 20:11:26 | spatel | Let me check.. | |
| 20:12:58 | spatel | I found this lines in nova-compute - AMQP server on 192.168.1.11:5672 is unreachable: timed out. Trying again in 0 seconds.: socket.timeout: timed out | |
| 20:13:11 | spatel | Look like issue is related to rabbit.. hmm | |
| 20:13:22 | spatel | but cluster status is green | |
| 20:13:56 | spatel | Let me destroy rabbit and rebuild it to see if it come clean | |
| 20:27:15 | spatel | dansmith look like it was rabbit issue, after re-building rabbit I can see correct count on hypervisor stats :) | |
| 20:28:00 | dansmith | spatel: good, that's why I was recommending you not just fix it manually because it's an indication of something else | |
| 20:28:48 | spatel | Thank you for staying with me :) | |
| 20:28:56 | spatel | I was about to blowup DB.. haha | |
| 20:29:55 | spatel | dansmith I have last question, How do i tell nova to just limit number of VMs per kvm host? | |
| 20:30:30 | spatel | I have 3 nodes and just want to stick with 10 vms per compute nodes limit so i don't blow up again | |
| 20:31:14 | dansmith | I don't know that you can, easily.. you might be able to hack that up with placement, a custom resource class, and flavor extra specs, but it would be complicated | |
| 20:31:47 | dansmith | better to set memory overcommit to 1.0 reserved memory enough to run your host services and then it will limit to whatever will fit without creating too much memory pressure | |
| 20:32:30 | spatel | In my case i have controller and compute on same node. | |
| 20:32:32 | spatel | This is very small environment for small budget. | |
| 20:33:08 | spatel | I like the idea of memory overcommit to 1.0 | |
| 20:34:04 | spatel | I thought nova has config setting per compute node to tell number of VM allow to run. | |
| 20:34:50 | dansmith | not that I know of.. generally such a number would make no sense.. one 32G instance might fit where 16 2GB instances would fit.. "number of instances" is not a very useful number for most people | |
| 20:57:05 | opendevreview | Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585 | |
| #openstack-nova - 2023-05-06 | |||
| 08:55:44 | manuvakery1 | Hi while running nova online-data-migration I am getting following error | |
| 08:55:47 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage found, done = migration_meth(ctxt, count) | |
| 08:55:47 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage File "/usr/lib/python2.7/site-packages/nova/objects/request_spec.py", line 732, in migrate_instances_add_request_spec | |
| 08:55:47 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage [req-0d649e34-56ab-4f5e-8f3b-2a062097a702 - - - - -] Error attempting to run <function migrate_instances_add_request_spec at 0x7facf137a500>: AttributeError: 'NoneType' object has no attribute 'hosts' | |
| 08:55:47 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage Traceback (most recent call last): | |
| 08:55:47 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage File "/usr/lib/python2.7/site-packages/nova/cmd/manage.py", line 679, in _run_migration | |
| 08:55:48 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage _create_minimal_request_spec(context, instance) | |
| 08:55:48 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage File "/usr/lib/python2.7/site-packages/nova/objects/request_spec.py", line 705, in _create_minimal_request_spec | |
| 08:55:50 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage scheduler_utils.setup_instance_group(context, request_spec) | |
| 08:55:50 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage File "/usr/lib/python2.7/site-packages/nova/scheduler/utils.py", line 890, in setup_instance_group | |
| 08:55:52 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage request_spec.instance_group.hosts = list(group_info.hosts) | |
| 08:55:52 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage AttributeError: 'NoneType' object has no attribute 'hosts' | |