Earlier  
Posted Nick Remark
#openstack-nova - 2021-01-13
13:31:10 bauzas brinzhang: ack, will look at your replies then
14:44:02 openstackgerrit Takashi Natsume proposed openstack/python-novaclient master: Deprecate agent commands and APIs https://review.opendev.org/c/openstack/python-novaclient/+/769068
16:19:12 melwitt stephenfin, artom: would like to have your numa expert review on this patch (and the func test below it) please if you could spare some time this week https://review.opendev.org/c/openstack/nova/+/769614
16:19:50 artom melwitt, will take a look tomorrow
16:20:03 melwitt thanks!
16:20:20 artom Hopefully before then, actually
16:29:51 openstackgerrit Lee Yarwood proposed openstack/nova master: WIP libvirt: Record the machine_type of instances in system_metadata https://review.opendev.org/c/openstack/nova/+/767533
16:29:55 openstackgerrit Lee Yarwood proposed openstack/nova master: WIP nova-manage: Add commands for managing instance machine type https://review.opendev.org/c/openstack/nova/+/769548
16:29:57 openstackgerrit Lee Yarwood proposed openstack/nova master: WIP nova-status: Add hw_machine_type check for libvirt instances https://review.opendev.org/c/openstack/nova/+/770643
16:56:44 sean-k-mooney artom: i havenet made any changes yet but i responed to your questions in https://review.opendev.org/c/openstack/nova-specs/+/764999/2/specs/wallaby/approved/libvirt-vdpa-support.rst can you re review
16:57:09 sean-k-mooney i have some WIP patches up as well https://review.opendev.org/q/topic:%22vhost-vdpa%22+(status:open%20OR%20status:merged)
16:57:27 openstackgerrit Balazs Gibizer proposed openstack/nova master: DNM try to replace retry with libvirt event in detach https://review.opendev.org/c/openstack/nova/+/770246
16:57:45 sean-k-mooney its not fully complete as i still need to extend the pci tracker and where we create the pci request
16:58:01 sean-k-mooney and i need to test it locally and write tests and docs
16:58:28 sean-k-mooney but it has the outline of about 2 thirds of the code
16:59:26 artom sean-k-mooney, cool - I want to un-WIP my socket affinity spec, then I'll take a look
17:10:04 markguz_ Hi nova folks. THe good people at #openstack-ironic thought it might be a good idea for me to ask about my problem here.
17:12:04 markguz_ I've got an issue where my ironic instance spawning is getting stuck for +/- 10mins just after i start deployment. The nova-scheduler picks the correct compute node and then the nova-compute node reports "Starting instance... _do_build_and_run_instance"
17:12:40 markguz_ then nothing happens for about 10 mins. then suddenly the process starts and deployment continues..
17:13:30 markguz_ TheJulia over at ironic thinks that the process is getting stuck at the scheduling stage. The vm reports Building, but the task status sits at "none" during that 10mins
17:14:43 markguz_ For VMs there is no delay, only for BMs. This is Rocky. and it is not a busy deployment. Very little activity going. We can go days without spawning bm or vm instances
17:15:35 markguz_ I've been trying to dig around the code to see what happens when the compute node reports "Starting instance... _do_build_and_run_instance" but it's very hard to follow
17:16:43 markguz_ If anyone could help me follow the rabbit through the rabbit hole and see exactly what is happening, I'd be most appreciative
17:23:58 melwitt markguz_: if it's landed on the compute node already, we don't consider that to be "scheduling" as it's already been scheduled/placed. but to dig into this further you'll want to trace the request id of the line that says "Starting instance" in the nova-compute log and see if you can see where it stops making progress. that would help
17:25:31 sean-k-mooney if this is ironic by they way whe n we get do do_build_and_run_instance at some point the compute manager will hand of to the ironic driver which will call ironic to provision the node
17:26:27 melwitt right. would want to look and see if he can verify it's gotten to that point
17:26:58 sean-k-mooney https://opendev.org/openstack/nova/src/branch/master/nova/compute/manager.py#L2186 starting to build instance is right at the top
17:27:31 sean-k-mooney the we save the task state at None and vmstate building
17:27:41 melwitt yeah, I know. just saying he can trace the request id to see how far it gets
17:27:54 melwitt after that
17:29:09 markguz_ here's a grep of the req id for an instance out of the logs http://paste.openstack.org/show/801601/
17:29:44 markguz_ nothing between 9.40 and 10.05
17:31:14 sean-k-mooney do you have the concurnet build limit set
17:31:15 stephenfin melwitt: Comments left
17:31:17 stephenfin sean-k-mooney: ^
17:31:20 sean-k-mooney for the compute service
17:31:26 sean-k-mooney it defaults to 10 i belive
17:31:40 TheJulia markguz_: something between scheduling and initial network setup :\
17:31:55 sean-k-mooney ock "compute_resources" acquired by "nova.compute.resource_tracker.instance_claim" :: waited 1495.269s
17:32:14 sean-k-mooney it looks like it was jsut waiting on the RT lock
17:32:20 melwitt yep, waited 24 min for the lock
17:32:47 melwitt thank you stephenfin
17:33:29 markguz_ why would the lock take 25mins?
17:34:15 sean-k-mooney the comput service is likely bussy starting up
17:34:24 sean-k-mooney you mentioned this is only after the iniall start right
17:34:41 sean-k-mooney or is this for each spawn
17:34:59 lyarwood just grep for the compute_resources lock and see what was holding it before?
17:35:35 sean-k-mooney ya that too
17:35:45 sean-k-mooney you could see how many other instance got the lock in that interval
17:36:22 markguz_ sean-k-mooney: it's every spawn. usually 10mins, sometimes longer and sometimes shorter
17:36:55 lyarwood does the resource tracker make external API calls in the Ironic driver?
17:38:46 sean-k-mooney the RT is shared so i dont think so
17:39:47 lyarwood yeah sorry I mean the code that refreshes it within the driver
17:40:11 sean-k-mooney im pretty sure this is the lock in question https://opendev.org/openstack/nova/src/branch/stable/rocky/nova/compute/manager.py#L2221-L2222
17:41:13 melwitt markguz_: I think you might be hitting https://bugs.launchpad.net/nova/+bug/1864122
17:41:15 openstack Launchpad bug 1864122 in OpenStack Compute (nova) "Instances (bare metal) queue for 30-60 seconds when managing a large amount of Ironic nodes" [Medium,Fix released] - Assigned to Jason Anderson (jasonandersonatuchicago)
17:41:50 sean-k-mooney yep that was what grabed the lock https://opendev.org/openstack/nova/src/branch/stable/rocky/nova/compute/resource_tracker.py#L159-L160
17:42:13 markguz_ i have +/- 240 nodes
17:42:24 markguz_ is that a large amount?
17:42:37 sean-k-mooney melwitt: yep a race on the lock with the update periodic task seams likely
17:42:41 melwitt if you read the bug it says can be seen around > 100 nodes
17:43:12 sean-k-mooney markguz_: how many ironic compute services do you have
17:43:15 melwitt that fix is available in ussuri and onward, it was not backported because it requires a newer version of oslo.concurrency
17:43:24 markguz_ sean-k-mooney: 1
17:43:29 sean-k-mooney i belive the periodic will only update the resouce usage for the nodes that are assgined to it
17:44:10 sean-k-mooney so i think you can scale it by deploying more compute service instances TheJulia is that correct?
17:45:03 sean-k-mooney markguz_: if you have 3 contolers i would suggest running an ironic nova compute service instance on each assuming that makes sense to TheJulia or others
17:45:07 TheJulia sean-k-mooney: yes, you can, you should just be able to run multiple instances
17:45:19 TheJulia markguz_: ^^^ instances of nova-compute configured for ironic
17:45:33 stephenfin melwitt: comments left on the bug report too
17:46:09 sean-k-mooney markguz_: the other thing you could do is reduce the interval of the periodic
17:46:21 melwitt hm, I thought you needed to configure node partitioning to do that
17:46:23 sean-k-mooney we fixed it by chanigin the type of lock we use
17:46:46 melwitt "conductor groups"
17:46:52 TheJulia melwitt: only to force specific grouping/allocation into specific grouping
17:46:56 sean-k-mooney that on the ironic side i think
17:47:01 melwitt it's not
17:47:05 TheJulia its on both sides
17:47:12 sean-k-mooney ah ok
17:47:19 markguz_ peridoc_task_interval is set to 240
17:47:19 melwitt well, it might be but you have to do it on the nova side too
17:47:24 TheJulia otherwise it runs a hash ring based upon the node list
17:47:31 TheJulia and the group is just a key in the hash ring
17:48:01 sean-k-mooney we improved this in nova by using oslos fair locks
17:48:02 TheJulia the nova side name is a little different because naming_is_fun^TM
17:48:02 sean-k-mooney https://review.opendev.org/c/openstack/nova/+/711528/2
17:48:23 markguz_ so we use this in a lab env and when we spin up baremetal we need to spin up a spcific node as they are connected to specific hardware that is being tested
17:48:39 sean-k-mooney but that was only done in ussuri
17:48:48 TheJulia sean-k-mooney: ohhhhh neat
17:49:02 sean-k-mooney so we would have to backport it unfortunetlly im not sure oslo has the required support in rocky let me check
17:49:13 melwitt ok conductor groups are not available until stein anyways
17:49:30 sean-k-mooney we would need oslo.concurrancy 3.29.0 to backport it
17:49:39 markguz_ if i run multiple computes i'm guessing that i will need to change how i call a instance. right now i use the "avail_zone:compute_host:bm_uuid" trick
17:49:40 melwitt yeah, I said all of that earlier
17:49:55 sean-k-mooney stable rocky is oslo.concurrency===3.27.0
17:50:07 TheJulia markguz_: yeah, :\
17:50:14 melwitt yes, the patch that added fair locks bumped the oslo.concurrency version
17:50:22 sean-k-mooney so it can go back to stien
17:50:32 sean-k-mooney but not rocky
17:50:40 melwitt so it wasn't bumped until ussuri

Earlier   Later