| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-02-16 | |||
| 17:08:01 | mriedem | melwitt: not necessarily a new volume | |
| 17:08:11 | cfriesen | mriedem: I think previously you would have an instance in ERROR state with an attached volume. I think you could have done a "rebuild" on it from the error state. | |
| 17:08:15 | mriedem | melwitt: i can create a volume in cinder, and use it for bfv, and specify delete_on_termination | |
| 17:08:19 | mriedem | i'm not sure why you'd do that though | |
| 17:08:51 | mriedem | cfriesen: i thought you were asking about instance.get_network_info() | |
| 17:08:58 | mriedem | not self.network_api.get_instance_nw_info(context, instance) | |
| 17:09:03 | mriedem | self.network_api.get_instance_nw_info(context, instance) will rebuild the info cache | |
| 17:09:07 | melwitt | mriedem: if delete_on_termination was set and it failed the build and you had to delete the instance and start again, you would have to create a new volume, right? | |
| 17:09:14 | mriedem | instance.get_network_info() is the same as instance.info_cache.network_info | |
| 17:09:26 | mriedem | melwitt: yes | |
| 17:10:17 | melwitt | that's what I was saying to cfriesen, you would only have to create a new volume if you had used delete_on_termination. else the volume would be only detached and you could use it again with a new instance | |
| 17:10:38 | smcginnis | You might want a bfv deleted on instance delete if you are using it just to manage your storage capacity separate from your local n-compute local storage. | |
| 17:11:07 | cfriesen | mriedem: how do I know when I can call instance.get_network_info()? Specifically I'm looking at nova.compute.rpcapi.pre_live_migration()....I want to adjust the RPC timeout based on the number of network ports. | |
| 17:11:12 | mriedem | smcginnis: so not the root disk | |
| 17:11:15 | mriedem | application data | |
| 17:11:30 | mriedem | i just figure people that create a volume directly in cinder and use it to bfv care about re-using the volume, | |
| 17:11:36 | mriedem | and people that let nova create the volume for you, don't care | |
| 17:11:47 | smcginnis | Eh, less useful maybe just for application data, but still can be used in that way. | |
| 17:12:00 | mriedem | cfriesen: you can call it at any time...? | |
| 17:12:13 | smcginnis | mriedem: I would think that is usually the case that they do care about that data though if they create in cinder first. | |
| 17:12:23 | mriedem | smcginnis: yeah | |
| 17:12:57 | mriedem | cfriesen: presumably we rebuild the nw info cache before starting live migration anyway | |
| 17:13:21 | mriedem | cfriesen: yes we do | |
| 17:13:22 | mriedem | network_info = self.network_api.get_instance_nw_info(context, instance) | |
| 17:13:28 | mriedem | in pre_live_migration in the ComputeManager | |
| 17:13:30 | cfriesen | mriedem: I'm confused (clearly). in the existing code, in some places where they want network_info they look at instance.info_cache.network_info, and in other places they call | |
| 17:13:32 | mriedem | that will refresh the nw info cache | |
| 17:13:32 | cfriesen | self.network_api.get_instance_nw_info( | |
| 17:13:33 | cfriesen | context, instance) | |
| 17:13:38 | cfriesen | whoops, bad paste, sorry | |
| 17:13:47 | mriedem | so at that point, the instance.info_cache.network_info is up to date | |
| 17:13:55 | mriedem | so if you're going to use it to count ports to adjust rpc call timeout, you should be good | |
| 17:15:13 | cfriesen | no, I need to have the ports in the rpc code that is *calling* pre_live_migration on the dest node. | |
| 17:15:23 | cfriesen | so updating in pre_live_migration doesn't help | |
| 17:15:36 | mriedem | let me dig up the live migration call chart | |
| 17:15:45 | mriedem | of which i still owe fried_rice a shiny nickel | |
| 17:16:01 | fried_rice | \o/ | |
| 17:16:02 | cfriesen | I think live_migration() on the source node calls pre_live_migration() on the dest | |
| 17:16:39 | cfriesen | or rather _do_live_migration() I guess | |
| 17:17:16 | mriedem | f me if i can ever find anything in our docs | |
| 17:17:26 | mriedem | https://docs.openstack.org/nova/latest/reference/live-migration.html | |
| 17:17:59 | mriedem | cfriesen: yeah you're right | |
| 17:18:07 | openstackgerrit | Matthew Edmonds proposed openstack/nova-specs master: PowerVM Virt Integration (Rocky) https://review.openstack.org/545111 | |
| 17:18:34 | mriedem | and _do_live_migration doesn't refresh the instance nw info cache before calling pre_live_migration | |
| 17:18:50 | cfriesen | okay, so I probably need to call self.network_api.get_instance_nw_info(context, instance) ? | |
| 17:18:57 | mriedem | cfriesen: so from _do_live_migration if you call instance.get_network_info() then you're just getting whatever is current for the instance in the db | |
| 17:19:06 | mriedem | if you really want to refresh to get the latest | |
| 17:19:27 | cfriesen | what would cause that to be stale? if someone was in the middle of attaching an interface or something? | |
| 17:19:31 | mriedem | but we have the network events from neutron if that changes, and the heal_instance_info_cache periodic | |
| 17:19:44 | mriedem | umm | |
| 17:20:10 | mriedem | yeah, | |
| 17:20:23 | mriedem | since we don't set a task_state for attaching/detaching an interface, you can live migrate the server at the same time | |
| 17:20:36 | mriedem | if the instance.task_state was not None (attaching/detaching), you couldn't do a live migration | |
| 17:20:58 | mriedem | thta's something i've always wondered about, | |
| 17:21:09 | mriedem | why we don't change task_state while attaching/detaching interfaces and volumes | |
| 17:22:30 | cfriesen | might be a fun stress test. ping-pong live migrations while attaching/detaching interfaces/volumes | |
| 17:24:46 | mriedem | fun like a hernia | |
| 17:26:51 | mriedem | cfriesen: so you're probably better off just starting small and getting the current info cache in _do_live_migration and basing the rpc timeout on that | |
| 17:36:11 | cfriesen | mriedem: agreed, it should be good enough. Is this idea of varying the RPC timeout based on number of attached interfaces something that you think would be generally useful? We're seeing it take ~1.5 sec per interface in pre_live_migration() to update the ports, but I'm not sure if that's typical or due to our neutron changes. | |
| 17:36:57 | superdan | I don't really like the idea of doing that | |
| 17:37:16 | cfriesen | and in nova those calls to neutron are serialized, so with a large number of interfaces it's possible to eat up a good chunk of the RPC timeout just doing the port updates | |
| 17:37:17 | superdan | I'd rather see something in olso.messaging that heartbeats running calls so we have a soft and hard timeout range | |
| 17:37:17 | mriedem | i seem to remember having similar conversations for other rpc calls that can take a long time, but can't remember details, | |
| 17:37:35 | mriedem | or if they were just "should we add a specific rpc timeout config option for this *one* really bad operation?" | |
| 17:37:49 | superdan | or work on not chaining rpc calls together | |
| 17:38:06 | mriedem | cfriesen: and this is with https://review.openstack.org/#/c/465787/ applied right? | |
| 17:38:25 | cfriesen | we internally already hack the pre_live_migration timeout for block-live-migration to allow time to download the image from glance. | |
| 17:38:31 | mriedem | cfriesen: do you happen to have any rough numbers on how much ^ helps with live migration with >1 ports? | |
| 17:39:03 | mriedem | are you using the image cache? | |
| 17:39:17 | cfriesen | mriedem: yes, it's with that applied. let me see if I can dig something up. | |
| 17:39:51 | cfriesen | mriedem: I think so, but that still means it could hit the first instance that wants that image. | |
| 17:41:01 | mriedem | on the dest host | |
| 17:41:09 | mriedem | the image would be cached on the source host but that doesn't help you | |
| 17:41:30 | mriedem | you guys don't use ceph? | |
| 17:41:37 | cfriesen | mriedem: up to the end user | |
| 17:41:48 | mriedem | ok; figured with all of the live migration you seem to do, you'd push ceph | |
| 17:41:58 | cfriesen | mriedem: we support compute nodes with ceph, with local qcow2, and with local thin LVM | |
| 17:42:45 | cfriesen | mriedem: some installs are really small, like 2 all-in-one nodes | |
| 17:48:55 | cfriesen | mriedem: looking at my notes, that patch cut the neutron load by a factor of 3 and reduced lock contention in nova. We also had an oslo.lockutils change to introduce fair locks--they merged it but then it seemed to cause mysterious issues in the CI for a couple of projects so they reverted it. | |
| 17:49:39 | mriedem | how many ports on that instance? | |
| 17:49:41 | mriedem | in that test? | |
| 17:50:16 | cfriesen | actually, wait, that factor of 3 reduction was for removing redundant calls for the same instance | |
| 17:51:46 | cfriesen | 16 ports on the instance (our max) | |
| 17:57:09 | cfriesen | mriedem: found some numbers. as of april last year each call to get_instance_nw_info() cost roughly 200ms plus about 125ms per port. And without that patch there are O(N) network-changed events in a live-migration (where N is the number of interfaces), so the overall cost is O(N^2) | |
| 17:57:17 | cfriesen | with that patch, it drops to O(N) | |
| 17:58:37 | mriedem | cool | |
| 17:58:48 | mriedem_lunch | i'll feast on those results at lunch | |
| 18:30:02 | openstackgerrit | Ken'ichi Ohmichi proposed openstack/nova master: Trivial: Update help of enabled_filters https://review.openstack.org/545431 | |
| 18:31:23 | openstackgerrit | Ken'ichi Ohmichi proposed openstack/nova master: Trivial: Update help of enabled_filters https://review.openstack.org/545431 | |
| 19:06:06 | mrjk_ | mriedem, ok, thank you for your hints | |
| 19:08:17 | mnaser | if nova creates neutron ports when it boots an instance, and the port is manually detached with nova instance-detach, the port is deleted (as per preserve_on_delete=False).. anyone else feel this behaviour is bleh? | |
| 19:08:29 | mnaser | if i detach a port manually, it means i want to use it | |
| 19:15:11 | imacdonn | hmm, interesting .... it could be argued that if you want to make your ports available for use with other instances, you should pre-create them separately | |
| 19:18:22 | imacdonn | assuming that, perhaps you shouldn't be allowed to detach a port that was created along with the instance (thinking out loud) | |
| 19:29:18 | TheJulia | jroll: do you, or anyone else, remember if there was any outcome from the duplicate placement issues we encountered yesterday with ironic's multinode grenade job on queens->master ? It passed without issues once, revised the patch slightly and hit the same issue again | |
| 19:29:53 | jroll | TheJulia: I didn't see anything, fried_rice said he might look into it | |
| 19:30:03 | jroll | I probably won't get bandwidth to investigate today | |
| 19:30:23 | jroll | all I got was that yes, we likely triggered a rebalance of ironic n-cpu things | |
| 19:36:39 | mrjk_ | cfriesen, I'm reading back your comments, "I'd suggest keeping the number of servers in a single boot request small enough" => Google does not answer to this question :) | |
| 19:37:30 | mrjk_ | That means I don't have the control on this issue ? (I was looking for a max cap allowed instance per requests) | |