Earlier  
Posted Nick Remark
#openstack-nova - 2023-05-05
19:47:05 dansmith you need to look at compute_nodes.running_vms to see which one is still reporting instances
19:47:31 spatel let me take a look at that table
19:51:06 spatel I am not able to find that tables in DB
19:51:25 spatel it should be inside nova/instance correct?
19:52:07 dansmith instance (assuming you meant instances) is a table, compute_nodes is a table, running_vms is a column in the compute_nodes table
19:53:23 spatel found it
19:56:00 spatel Yes, I can see them there that on node1 - https://paste.opendev.org/show/bnhl6ZeHXIA3Tjk1ucV6/
19:56:51 spatel do you think just update those number in table is enough?
19:56:55 dansmith that's three nodes
19:56:57 dansmith no
19:57:05 dansmith I mean, that will make the number change, but it's not the right fix
19:57:17 dansmith you need to select the hostname along with the count to know which is which
19:57:27 dansmith nova-compute should be updating those numbers
19:58:19 dansmith select host,hypervisor_hostname,running_vms from compute_nodes;
19:58:48 spatel https://paste.opendev.org/show/bwohGVNFtteg3J4x5qUY/
19:59:10 spatel at present on ctrl node there are no VM running..
19:59:22 spatel at present on ctrl1 and ctrl3 node there are no VM running..
19:59:37 spatel That entry should be zero technically
19:59:50 dansmith is nova-compute running on each of those three nodes?
19:59:59 dansmith because it should be updating that number every few minutes
20:00:03 spatel yes its running
20:00:30 spatel all services showing fine.. I have restarted them
20:00:32 spatel no nasty logs or errors anywhere
20:02:50 dansmith they should all be iterating over their instances regularly and updating those numbers
20:05:12 dansmith perhaps it's not doing that if there are no instances (although you said there was one, so at least that one should be correct)
20:06:58 spatel out of 3 nodes only node2 has 1 VM running and rest are empty
20:07:41 dansmith yeah, so that node should show 1 in the database and doesn't, which to me means something is wrong (unless it hasn't run update_available_resource yet)
20:07:43 spatel Thinking to reboot all 3 nodes to start with fresh troubleshooting
20:08:30 spatel who will run update_available_resource task? compute nodes correct?
20:08:53 dansmith nova-compute does it
20:09:11 spatel may be rabbitMQ is in zombie state... I have checked cluster_status and its showing all good but who knows..
20:09:47 dansmith there should be errors in nova-compute if so, but hard to say after something like an oom
20:11:26 spatel Let me check..
20:12:58 spatel I found this lines in nova-compute - AMQP server on 192.168.1.11:5672 is unreachable: timed out. Trying again in 0 seconds.: socket.timeout: timed out
20:13:11 spatel Look like issue is related to rabbit.. hmm
20:13:22 spatel but cluster status is green
20:13:56 spatel Let me destroy rabbit and rebuild it to see if it come clean
20:27:15 spatel dansmith look like it was rabbit issue, after re-building rabbit I can see correct count on hypervisor stats :)
20:28:00 dansmith spatel: good, that's why I was recommending you not just fix it manually because it's an indication of something else
20:28:48 spatel Thank you for staying with me :)
20:28:56 spatel I was about to blowup DB.. haha
20:29:55 spatel dansmith I have last question, How do i tell nova to just limit number of VMs per kvm host?
20:30:30 spatel I have 3 nodes and just want to stick with 10 vms per compute nodes limit so i don't blow up again
20:31:14 dansmith I don't know that you can, easily.. you might be able to hack that up with placement, a custom resource class, and flavor extra specs, but it would be complicated
20:31:47 dansmith better to set memory overcommit to 1.0 reserved memory enough to run your host services and then it will limit to whatever will fit without creating too much memory pressure
20:32:30 spatel In my case i have controller and compute on same node.
20:32:32 spatel This is very small environment for small budget.
20:33:08 spatel I like the idea of memory overcommit to 1.0
20:34:04 spatel I thought nova has config setting per compute node to tell number of VM allow to run.
20:34:50 dansmith not that I know of.. generally such a number would make no sense.. one 32G instance might fit where 16 2GB instances would fit.. "number of instances" is not a very useful number for most people
20:57:05 opendevreview Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585
#openstack-nova - 2023-05-06
08:55:44 manuvakery1 Hi while running nova online-data-migration I am getting following error
08:55:47 manuvakery1 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage found, done = migration_meth(ctxt, count)
08:55:47 manuvakery1 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage File "/usr/lib/python2.7/site-packages/nova/objects/request_spec.py", line 732, in migrate_instances_add_request_spec
08:55:47 manuvakery1 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage [req-0d649e34-56ab-4f5e-8f3b-2a062097a702 - - - - -] Error attempting to run <function migrate_instances_add_request_spec at 0x7facf137a500>: AttributeError: 'NoneType' object has no attribute 'hosts'
08:55:47 manuvakery1 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage Traceback (most recent call last):
08:55:47 manuvakery1 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage File "/usr/lib/python2.7/site-packages/nova/cmd/manage.py", line 679, in _run_migration
08:55:48 manuvakery1 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage _create_minimal_request_spec(context, instance)
08:55:48 manuvakery1 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage File "/usr/lib/python2.7/site-packages/nova/objects/request_spec.py", line 705, in _create_minimal_request_spec
08:55:50 manuvakery1 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage scheduler_utils.setup_instance_group(context, request_spec)
08:55:50 manuvakery1 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage File "/usr/lib/python2.7/site-packages/nova/scheduler/utils.py", line 890, in setup_instance_group
08:55:52 manuvakery1 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage request_spec.instance_group.hosts = list(group_info.hosts)
08:55:52 manuvakery1 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage AttributeError: 'NoneType' object has no attribute 'hosts'
08:55:54 manuvakery1 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage
08:55:54 manuvakery1 2023-05-06 07:54:26.777 256132 ERROR nova.objects.keypair [req-0d649e34-56ab-4f5e-8f3b-2a062097a702 - - - - -] Some instances are still missing keypair information. Unable to run keypair migration at this time.
08:56:10 manuvakery1 what am I missing here
09:54:27 manuvakery1 Ignore it worked after deleting some broken instances
#openstack-nova - 2023-05-07
01:11:45 opendevreview Takashi Natsume proposed openstack/placement master: Fix a wrong assertion method https://review.opendev.org/c/openstack/placement/+/861489
01:12:15 opendevreview Takashi Natsume proposed openstack/nova master: Update contributor guide for 2023.2 Bobcat https://review.opendev.org/c/openstack/nova/+/876447
#openstack-nova - 2023-05-08
04:58:43 opendevreview Amit Uniyal proposed openstack/nova master: WIP: Delete dangling bdms https://review.opendev.org/c/openstack/nova/+/882284
08:56:51 opendevreview Balazs Gibizer proposed openstack/nova stable/2023.1: Reproduce asym NUMA mixed CPU policy bug https://review.opendev.org/c/openstack/nova/+/882508
08:56:52 opendevreview Balazs Gibizer proposed openstack/nova stable/2023.1: Handle zero pinned CPU in a cell with mixed policy https://review.opendev.org/c/openstack/nova/+/882509
08:57:59 opendevreview Balazs Gibizer proposed openstack/nova stable/zed: Reproduce asym NUMA mixed CPU policy bug https://review.opendev.org/c/openstack/nova/+/882530
08:58:00 opendevreview Balazs Gibizer proposed openstack/nova stable/zed: Handle zero pinned CPU in a cell with mixed policy https://review.opendev.org/c/openstack/nova/+/882531
08:59:16 opendevreview Balazs Gibizer proposed openstack/nova stable/yoga: Reproduce asym NUMA mixed CPU policy bug https://review.opendev.org/c/openstack/nova/+/882532
08:59:17 opendevreview Balazs Gibizer proposed openstack/nova stable/yoga: Handle zero pinned CPU in a cell with mixed policy https://review.opendev.org/c/openstack/nova/+/882533
09:00:07 opendevreview Balazs Gibizer proposed openstack/nova stable/xena: Reproduce asym NUMA mixed CPU policy bug https://review.opendev.org/c/openstack/nova/+/882534
09:00:08 opendevreview Balazs Gibizer proposed openstack/nova stable/xena: Handle zero pinned CPU in a cell with mixed policy https://review.opendev.org/c/openstack/nova/+/882535
09:01:42 opendevreview Balazs Gibizer proposed openstack/nova stable/wallaby: Reproduce asym NUMA mixed CPU policy bug https://review.opendev.org/c/openstack/nova/+/882536
09:01:42 opendevreview Balazs Gibizer proposed openstack/nova stable/wallaby: Handle zero pinned CPU in a cell with mixed policy https://review.opendev.org/c/openstack/nova/+/882537
09:03:51 opendevreview Danylo Vodopianov proposed openstack/nova master: Napatech SmartNIC support https://review.opendev.org/c/openstack/nova/+/859577
09:44:12 kashyap dansmith: When you're about - I added "Affects" "kernel-package (Ubuntu)" here: https://bugs.launchpad.net/ubuntu/+source/kernel-package/+bug/2018612
09:44:50 kashyap Added some environment details, and I also just asked on #ubunutu-kernel to make sure it's the right package name.
10:03:46 opendevreview Amit Uniyal proposed openstack/nova master: WIP: Delete dangling bdms https://review.opendev.org/c/openstack/nova/+/882284
11:55:57 kashyap dansmith: Folks from #ubuntu-kernel say that the kernel in the CirrOS 0.5.2 image is unsupported.
11:56:38 kashyap The kernel ("5.3.0-26-generic") from CirrOS 0.5.2 _used_ to be LTS; but right now it's 5.4 (from Focal).
11:59:00 kashyap I'll ask the person responding to also respond on the bug
12:11:37 kashyap dansmith: So that was the current diagnosis from #ubuntu-kernel folks:
12:11:43 kashyap CirrOS 0.5.1 was built in march 2020, before Focal release, which can explain the 5.3 kernel [that is crashing]. And CirrOS 0.5.2 didn't upgrade the kernel.
12:18:24 kgube Hi bauzas, could you have another look at this? https://review.opendev.org/c/openstack/nova-specs/+/877233
12:28:29 kashyap dansmith: I'm filing a CirrOS 0.5.2 issue to refresh the image and update the packages.
12:30:03 ykarel just in case there is an image cirros 0.6.0/0.6.1 if that helps
12:30:45 kashyap ykarel: Yeah; I know it's there. But we're trying to tackle a specific issue w/ 0.5.2 :)
12:31:11 ykarel ack :)
12:42:53 opendevreview Merged openstack/placement master: Fix a wrong assertion method https://review.opendev.org/c/openstack/placement/+/861489
12:48:27 kashyap dansmith: And that's the CirrOS issue I filed: https://github.com/cirros-dev/cirros/issues/102
12:49:41 kashyap (More details in the Nova bug you filed)
13:04:53 kashyap sean-k-mooney[m]: As promised, the Ubunutu folks picked up my device-detach backport request; and 'paelzer' (Christian Ehrhardt) filed a tracker on my behalf: https://bugs.launchpad.net/ubuntu/+source/qemu/+bug/2018733/
13:16:33 dansmith kashyap: ack, cool, I know we have a todo to upgrade to 6.x but I think there are other factors at play
13:46:43 dansmith gouthamr: I created this https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/882483 to try to convert our live migration multinode job over to cephadm, but it looks like the two nodes aren't setup properly

Earlier   Later