| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-05-05 | |||
| 19:59:10 | spatel | at present on ctrl node there are no VM running.. | |
| 19:59:22 | spatel | at present on ctrl1 and ctrl3 node there are no VM running.. | |
| 19:59:37 | spatel | That entry should be zero technically | |
| 19:59:50 | dansmith | is nova-compute running on each of those three nodes? | |
| 19:59:59 | dansmith | because it should be updating that number every few minutes | |
| 20:00:03 | spatel | yes its running | |
| 20:00:30 | spatel | all services showing fine.. I have restarted them | |
| 20:00:32 | spatel | no nasty logs or errors anywhere | |
| 20:02:50 | dansmith | they should all be iterating over their instances regularly and updating those numbers | |
| 20:05:12 | dansmith | perhaps it's not doing that if there are no instances (although you said there was one, so at least that one should be correct) | |
| 20:06:58 | spatel | out of 3 nodes only node2 has 1 VM running and rest are empty | |
| 20:07:41 | dansmith | yeah, so that node should show 1 in the database and doesn't, which to me means something is wrong (unless it hasn't run update_available_resource yet) | |
| 20:07:43 | spatel | Thinking to reboot all 3 nodes to start with fresh troubleshooting | |
| 20:08:30 | spatel | who will run update_available_resource task? compute nodes correct? | |
| 20:08:53 | dansmith | nova-compute does it | |
| 20:09:11 | spatel | may be rabbitMQ is in zombie state... I have checked cluster_status and its showing all good but who knows.. | |
| 20:09:47 | dansmith | there should be errors in nova-compute if so, but hard to say after something like an oom | |
| 20:11:26 | spatel | Let me check.. | |
| 20:12:58 | spatel | I found this lines in nova-compute - AMQP server on 192.168.1.11:5672 is unreachable: timed out. Trying again in 0 seconds.: socket.timeout: timed out | |
| 20:13:11 | spatel | Look like issue is related to rabbit.. hmm | |
| 20:13:22 | spatel | but cluster status is green | |
| 20:13:56 | spatel | Let me destroy rabbit and rebuild it to see if it come clean | |
| 20:27:15 | spatel | dansmith look like it was rabbit issue, after re-building rabbit I can see correct count on hypervisor stats :) | |
| 20:28:00 | dansmith | spatel: good, that's why I was recommending you not just fix it manually because it's an indication of something else | |
| 20:28:48 | spatel | Thank you for staying with me :) | |
| 20:28:56 | spatel | I was about to blowup DB.. haha | |
| 20:29:55 | spatel | dansmith I have last question, How do i tell nova to just limit number of VMs per kvm host? | |
| 20:30:30 | spatel | I have 3 nodes and just want to stick with 10 vms per compute nodes limit so i don't blow up again | |
| 20:31:14 | dansmith | I don't know that you can, easily.. you might be able to hack that up with placement, a custom resource class, and flavor extra specs, but it would be complicated | |
| 20:31:47 | dansmith | better to set memory overcommit to 1.0 reserved memory enough to run your host services and then it will limit to whatever will fit without creating too much memory pressure | |
| 20:32:30 | spatel | In my case i have controller and compute on same node. | |
| 20:32:32 | spatel | This is very small environment for small budget. | |
| 20:33:08 | spatel | I like the idea of memory overcommit to 1.0 | |
| 20:34:04 | spatel | I thought nova has config setting per compute node to tell number of VM allow to run. | |
| 20:34:50 | dansmith | not that I know of.. generally such a number would make no sense.. one 32G instance might fit where 16 2GB instances would fit.. "number of instances" is not a very useful number for most people | |
| 20:57:05 | opendevreview | Dan Smith proposed openstack/nova master: DNM: Test new ceph job configuration with nova https://review.opendev.org/c/openstack/nova/+/881585 | |
| #openstack-nova - 2023-05-06 | |||
| 08:55:44 | manuvakery1 | Hi while running nova online-data-migration I am getting following error | |
| 08:55:47 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage File "/usr/lib/python2.7/site-packages/nova/cmd/manage.py", line 679, in _run_migration | |
| 08:55:47 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage Traceback (most recent call last): | |
| 08:55:47 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage [req-0d649e34-56ab-4f5e-8f3b-2a062097a702 - - - - -] Error attempting to run <function migrate_instances_add_request_spec at 0x7facf137a500>: AttributeError: 'NoneType' object has no attribute 'hosts' | |
| 08:55:47 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage File "/usr/lib/python2.7/site-packages/nova/objects/request_spec.py", line 732, in migrate_instances_add_request_spec | |
| 08:55:47 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage found, done = migration_meth(ctxt, count) | |
| 08:55:48 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage File "/usr/lib/python2.7/site-packages/nova/objects/request_spec.py", line 705, in _create_minimal_request_spec | |
| 08:55:48 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage _create_minimal_request_spec(context, instance) | |
| 08:55:50 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage File "/usr/lib/python2.7/site-packages/nova/scheduler/utils.py", line 890, in setup_instance_group | |
| 08:55:50 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage scheduler_utils.setup_instance_group(context, request_spec) | |
| 08:55:52 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage AttributeError: 'NoneType' object has no attribute 'hosts' | |
| 08:55:52 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage request_spec.instance_group.hosts = list(group_info.hosts) | |
| 08:55:54 | manuvakery1 | 2023-05-06 07:54:26.777 256132 ERROR nova.objects.keypair [req-0d649e34-56ab-4f5e-8f3b-2a062097a702 - - - - -] Some instances are still missing keypair information. Unable to run keypair migration at this time. | |
| 08:55:54 | manuvakery1 | 2023-05-06 07:54:26.277 256132 ERROR nova.cmd.manage | |
| 08:56:10 | manuvakery1 | what am I missing here | |
| 09:54:27 | manuvakery1 | Ignore it worked after deleting some broken instances | |
| #openstack-nova - 2023-05-07 | |||
| 01:11:45 | opendevreview | Takashi Natsume proposed openstack/placement master: Fix a wrong assertion method https://review.opendev.org/c/openstack/placement/+/861489 | |
| 01:12:15 | opendevreview | Takashi Natsume proposed openstack/nova master: Update contributor guide for 2023.2 Bobcat https://review.opendev.org/c/openstack/nova/+/876447 | |
| #openstack-nova - 2023-05-08 | |||
| 04:58:43 | opendevreview | Amit Uniyal proposed openstack/nova master: WIP: Delete dangling bdms https://review.opendev.org/c/openstack/nova/+/882284 | |
| 08:56:51 | opendevreview | Balazs Gibizer proposed openstack/nova stable/2023.1: Reproduce asym NUMA mixed CPU policy bug https://review.opendev.org/c/openstack/nova/+/882508 | |
| 08:56:52 | opendevreview | Balazs Gibizer proposed openstack/nova stable/2023.1: Handle zero pinned CPU in a cell with mixed policy https://review.opendev.org/c/openstack/nova/+/882509 | |
| 08:57:59 | opendevreview | Balazs Gibizer proposed openstack/nova stable/zed: Reproduce asym NUMA mixed CPU policy bug https://review.opendev.org/c/openstack/nova/+/882530 | |
| 08:58:00 | opendevreview | Balazs Gibizer proposed openstack/nova stable/zed: Handle zero pinned CPU in a cell with mixed policy https://review.opendev.org/c/openstack/nova/+/882531 | |
| 08:59:16 | opendevreview | Balazs Gibizer proposed openstack/nova stable/yoga: Reproduce asym NUMA mixed CPU policy bug https://review.opendev.org/c/openstack/nova/+/882532 | |
| 08:59:17 | opendevreview | Balazs Gibizer proposed openstack/nova stable/yoga: Handle zero pinned CPU in a cell with mixed policy https://review.opendev.org/c/openstack/nova/+/882533 | |
| 09:00:07 | opendevreview | Balazs Gibizer proposed openstack/nova stable/xena: Reproduce asym NUMA mixed CPU policy bug https://review.opendev.org/c/openstack/nova/+/882534 | |
| 09:00:08 | opendevreview | Balazs Gibizer proposed openstack/nova stable/xena: Handle zero pinned CPU in a cell with mixed policy https://review.opendev.org/c/openstack/nova/+/882535 | |
| 09:01:42 | opendevreview | Balazs Gibizer proposed openstack/nova stable/wallaby: Handle zero pinned CPU in a cell with mixed policy https://review.opendev.org/c/openstack/nova/+/882537 | |
| 09:01:42 | opendevreview | Balazs Gibizer proposed openstack/nova stable/wallaby: Reproduce asym NUMA mixed CPU policy bug https://review.opendev.org/c/openstack/nova/+/882536 | |
| 09:03:51 | opendevreview | Danylo Vodopianov proposed openstack/nova master: Napatech SmartNIC support https://review.opendev.org/c/openstack/nova/+/859577 | |
| 09:44:12 | kashyap | dansmith: When you're about - I added "Affects" "kernel-package (Ubuntu)" here: https://bugs.launchpad.net/ubuntu/+source/kernel-package/+bug/2018612 | |
| 09:44:50 | kashyap | Added some environment details, and I also just asked on #ubunutu-kernel to make sure it's the right package name. | |
| 10:03:46 | opendevreview | Amit Uniyal proposed openstack/nova master: WIP: Delete dangling bdms https://review.opendev.org/c/openstack/nova/+/882284 | |
| 11:55:57 | kashyap | dansmith: Folks from #ubuntu-kernel say that the kernel in the CirrOS 0.5.2 image is unsupported. | |
| 11:56:38 | kashyap | The kernel ("5.3.0-26-generic") from CirrOS 0.5.2 _used_ to be LTS; but right now it's 5.4 (from Focal). | |
| 11:59:00 | kashyap | I'll ask the person responding to also respond on the bug | |
| 12:11:37 | kashyap | dansmith: So that was the current diagnosis from #ubuntu-kernel folks: | |
| 12:11:43 | kashyap | CirrOS 0.5.1 was built in march 2020, before Focal release, which can explain the 5.3 kernel [that is crashing]. And CirrOS 0.5.2 didn't upgrade the kernel. | |
| 12:18:24 | kgube | Hi bauzas, could you have another look at this? https://review.opendev.org/c/openstack/nova-specs/+/877233 | |
| 12:28:29 | kashyap | dansmith: I'm filing a CirrOS 0.5.2 issue to refresh the image and update the packages. | |
| 12:30:03 | ykarel | just in case there is an image cirros 0.6.0/0.6.1 if that helps | |
| 12:30:45 | kashyap | ykarel: Yeah; I know it's there. But we're trying to tackle a specific issue w/ 0.5.2 :) | |
| 12:31:11 | ykarel | ack :) | |
| 12:42:53 | opendevreview | Merged openstack/placement master: Fix a wrong assertion method https://review.opendev.org/c/openstack/placement/+/861489 | |
| 12:48:27 | kashyap | dansmith: And that's the CirrOS issue I filed: https://github.com/cirros-dev/cirros/issues/102 | |
| 12:49:41 | kashyap | (More details in the Nova bug you filed) | |
| 13:04:53 | kashyap | sean-k-mooney[m]: As promised, the Ubunutu folks picked up my device-detach backport request; and 'paelzer' (Christian Ehrhardt) filed a tracker on my behalf: https://bugs.launchpad.net/ubuntu/+source/qemu/+bug/2018733/ | |
| 13:16:33 | dansmith | kashyap: ack, cool, I know we have a todo to upgrade to 6.x but I think there are other factors at play | |
| 13:46:43 | dansmith | gouthamr: I created this https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/882483 to try to convert our live migration multinode job over to cephadm, but it looks like the two nodes aren't setup properly | |
| 13:47:08 | dansmith | gouthamr: that might require some help from a more ceph-literate person than me | |
| 13:48:49 | dansmith | they have different mon_host and fsid values, which I assume is .. not right | |
| 14:00:28 | kashyap | dansmith: Oh, good to know - if there's already an infra TODO to update to 0.6.x (CirrOS) | |
| 14:06:29 | dansmith | it was actually bauzas that had the patch up IIRC | |
| 14:10:38 | opendevreview | Danylo Vodopianov proposed openstack/nova-specs master: Add support for Napatech LinkVirt SmartNICs https://review.opendev.org/c/openstack/nova-specs/+/859290 | |
| 15:31:38 | gouthamr | hi dansmith: yes, "REMOTE_CEPH" doesn't work with cephadm; on multi-node you'd need that.. | |
| 15:32:04 | dansmith | gouthamr: ah, I see that's the other .. yeah, that patch | |
| 15:32:18 | dansmith | okay cool, well, we probably need to get that resolved before you drop the non-cephadm stuff | |
| 15:32:42 | gouthamr | yep | |
| 23:03:34 | opendevreview | Dan Smith proposed openstack/nova master: Populate ComputeNode.service_id https://review.opendev.org/c/openstack/nova/+/879904 | |
| 23:03:35 | opendevreview | Dan Smith proposed openstack/nova master: Add dest_compute_id to Migration object https://review.opendev.org/c/openstack/nova/+/879682 | |
| 23:03:35 | opendevreview | Dan Smith proposed openstack/nova master: Add compute_id columns to instances, migrations https://review.opendev.org/c/openstack/nova/+/879499 | |
| 23:03:36 | opendevreview | Dan Smith proposed openstack/nova master: Online migrate missing Instance.compute_id fields https://review.opendev.org/c/openstack/nova/+/879905 | |
| 23:03:36 | opendevreview | Dan Smith proposed openstack/nova master: Add compute_id to Instance object https://review.opendev.org/c/openstack/nova/+/879500 | |
| #openstack-nova - 2023-05-09 | |||
| 01:12:51 | opendevreview | sean mooney proposed openstack/os-vif master: [WIP] set default qos policy https://review.opendev.org/c/openstack/os-vif/+/881751 | |