Earlier  
Posted Nick Remark
#openstack-nova - 2018-04-24
16:33:05 melwitt (I don't know where reserve is currently happening)
16:33:24 TheJulia melwitt: in spawn when we begin to populate information in ironic about the instance
16:33:38 melwitt ohhh yeah. sorry
16:33:57 melwitt this network prep thing is called in compute/manager early
16:34:07 TheJulia no worries
16:34:08 TheJulia yeah
16:37:11 melwitt looking in the code for where reserve is currently happening and can't find it
16:37:31 jaypipes melwitt: yeah, we might want to move the network prep for block devices thing to being a method that is called from the virt driver and not the compute manager
16:38:07 jaypipes melwitt: so that the virt driver can dictate when it is appropriate to do that prep work
16:38:23 TheJulia melwitt: instance_uuid being set
16:38:23 dansmith melwitt: are you suggesting plugging vifs before we call virt spawn across the board?
16:38:37 jroll I did a poc for the pre-spawn method call TheJulia mentioned, fwiw: https://review.openstack.org/#/c/563722/
16:38:38 jaypipes dansmith: nope, not plugging vifs.
16:38:40 melwitt dansmith: no, I mean, where can we do the ironic node reserve call
16:39:04 dansmith jaypipes: okay I see vif plugging discussion above, but wasn't following
16:39:19 melwitt because ultimately we want to reserve the node before plugging vifs or attaching volumes. I just don't know where that's currently being done
16:39:28 TheJulia melwitt: let me get you the line, one moment
16:39:35 jroll melwitt: currently the ironic driver reserves the node within spawn()
16:39:46 jroll by sending instance_uuid in a PUT request
16:39:57 melwitt okay I see
16:39:57 jroll or node.update I guess, in ironicclient terms
16:40:25 melwitt I mean, the cheat would be move that call to "prepare networking etc" method
16:40:27 TheJulia https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L360
16:40:59 jroll melwitt: ya, that's the other version: https://review.openstack.org/#/c/563714/
16:41:05 TheJulia :)
16:41:07 jaypipes dansmith: no worries. we were discussing that ironic (actually, some cinder drivers) need to know the IP address of a port up front. and the prep_network_for-blokc_devices() virt driver API method was added in order to address that need. But that method does vif plugging (because that's unfortunately when a port's IP address is guaranteed to be set). We would like to see the IP address allocation decoupled from the vif plugging action.
16:41:22 jroll feels out of place to me but if the bug gets fixed then ¯\_(ツ)_/¯
16:41:26 melwitt jroll: ah, so that's what y'all mean by "lock"
16:41:44 dansmith jaypipes: okay
16:41:46 TheJulia s/lock/giant flag saying node is in use/
16:42:00 jroll melwitt: we also have internal locks which block certain actions on the node, so... sometimes :)
16:42:10 melwitt yeah, makes sense
16:42:18 melwitt heh
16:42:50 melwitt jaypipes, dansmith: but it can't be because DHCP, right?
16:43:01 dansmith melwitt: hmm?
16:43:26 dansmith melwitt: once you know the host, you can do the binding and get a port allocation at that point
16:43:36 melwitt dansmith: like, I had been thinking up to now that we get an IP allocated when we create the neutron port. but in the case of DHCP being used, that would not be true, right?
16:43:37 dansmith er, an address allocation for the port I mean
16:43:42 jaypipes melwitt: yes, when a port is in a subnet that is DHCP-enabled, the port does not get an IP address until after vif plugging (and the DHCP lease is done)
16:44:16 dansmith jaypipes: I think you get an address when you host-bind it, not exactly plug right?
16:44:23 jroll note that ironic doesn't bind the host to the port in nova, but later in ironic, because it can't be attached to the tenant network until it's deployed
16:45:01 melwitt dansmith: because ultimately what they need is to know the IP before they attach the volume. and so far, they have to plug the vif first in order to get the IP. and the problem that's happening is that the vif plug is happening outside of the node reserve lock, so things are racing
16:45:02 jaypipes dansmith: in the case of DHCP, the VIF needs to be fully set up and then the DHCP request made to the gateway, though, right? only after that will the port get an IP adderess.
16:45:28 dansmith jaypipes: you don't have to hit the dhcp server with a client to get an address
16:46:01 jaypipes dansmith: sorry, I wasn't aware that was possible.
16:46:08 dansmith melwitt: I'm not positive at which exact step neutron will assign an address I guess (plug vs. bind),
16:46:33 dansmith but I'm nearly positive it happens before the client is really set up,
16:46:47 melwitt dansmith: plug is just a local thing (os-vif in the case of libvirt) so it must be the bind, I think
16:46:47 dansmith otherwise you wouldn't be able to see what ip your instance is going to have until it has come up enough to have hit the dhcp server
16:46:50 dansmith which wouldn't make any sense
16:47:07 melwitt sean-k-mooney we need you!
16:47:10 melwitt :)
16:47:27 jroll I feel like IP allocation is done at port create time, but I would need to verify
16:47:35 TheJulia I'm 95% sure it is
16:47:38 dansmith jroll: exactly
16:47:43 melwitt that's what I said earlier
16:47:54 jroll dansmith: with or without a host binding, to be clear
16:47:57 melwitt so I was thinking, why not just ask neutron for the IP instead of doing plug_vifs
16:48:05 melwitt *plug_vifs early
16:48:15 dansmith jroll: I think it depends on how the network is setup whether you get it early or late, IIRC, but I dunno
16:48:25 jaypipes dansmith: by "port creation time" are you referring to when neutron port-create is done? because I'm pretty sure that *isn't* when IP allocation is done.
16:48:33 jroll could be, neutron is just a framework after all :)
16:48:49 dansmith I think it can happen at port-create time, and I think it can happen at host bind time
16:48:56 dansmith but I don't think it happens at vif_plug time
16:48:57 melwitt yes it's configurable https://developer.openstack.org/api-ref/network/v2/#ip-allocation-extension
16:49:27 dansmith think about when you're booting an instance and when you get to see the IP via the api in the scheme of it booting
16:49:29 dansmith usually you see it before it has even finished downloaded the image to the compute node right?
16:49:40 melwitt right, that's what I thought
16:49:58 melwitt but I thought is that only for static IPs and not DHCP?
16:50:04 melwitt I had thought it didn't matter
16:50:15 dansmith dhcp is just how you communicate the ip to the guest,
16:50:26 dansmith I don't think it changes how/when the port would get assigned an ip
16:50:51 jroll melwitt: so, ironic does a very late host binding of the port, because at that time it's put onto the tenant network (which we don't want during deployment). our plug_vifs call is the api endpoint for the code that does the host binding in ironic: https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L1620
16:51:35 melwitt yeah, good point. we use DHCP in the gate and yeah, pretty sure the IP is known once the neutron port is created. but we do know that's configurable, not necessarily true that the IP will be known at port create time depending on how the network is setup in neutron, seems like
16:52:08 jroll (and so during BFV deployment, prepare_networks_before_block_device_mapping is a fine time to put it on the tenant network, since we skip the deployment ramdisk)
16:52:18 melwitt jroll: yeah, figured it must be because plug_vifs was needed to get the IP
16:53:04 jroll melwitt: I assume it's just to hook up the networks, but not sure, I don't know this code well
16:53:09 artom Which one between flavor extra specs and image properties is arbitrary again?
16:53:17 dansmith artom: the former
16:53:30 dansmith melwitt: maybe mlavalle is around and could answer some questions
16:53:57 jroll ah yes, the IP comes from the BDM info: https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L399
16:53:57 artom dansmith, so operators can add a "foo" extra spec, and enable a filter (which one?) that would schedule all those foos to a certain aggregate?
16:54:12 jroll the plug_vifs call is just to get networking up
16:54:19 melwitt jroll: yeah, so back when we started talking about this today I was saying can we decouple the vif plugging from the IP get query and leave the vif plugging for after the node reserve like it used to be. not sure if the host bind should be behind reserve too though
16:54:54 dansmith artom: I have to go look for the linkage every time I'm asked.. I think there is an AggregateExtraSpecs filter or something that you use
16:54:57 jroll melwitt: the vif plugging is the host bind for us, but yeah, good question
16:55:16 artom dansmith, aha, thanks, I'll dig in the code then
16:55:19 melwitt jroll: this is the change that moved vif plug out from the node reserve in order to get an IP for the volume connector https://review.openstack.org/#/c/468353/19/nova/virt/ironic/driver.py
16:55:20 dansmith see I think in the late case, the host binding step is when you get your allocation
16:55:24 jroll melwitt: it feels like we could
16:56:06 melwitt dansmith: yeah. so would it be safe to do that outside of node reserve? jroll?
16:56:26 melwitt besides that, wouldn't that require a change to ironic API too?
16:56:26 jroll dansmith: melwitt: oh, right, that docstring tells us exactly that
16:56:54 dansmith melwitt: you can do host binding before spawn, that should be fine, you just can't do the plug before it
16:57:15 melwitt k
16:57:31 jroll melwitt: so, the problem we're seeing is when scheduling races with two instances to a node, they both try to do the plug because we haven't set that reservation. so I'm thinking if we do that reservation first thing, then we never hit this again
16:58:02 melwitt so it sounds like we have two options: decouple the host binding and do that in prepare_networks_before_block_device_mapping, get the IP for the volume attach, then reserve, then plug vifs etc
16:58:17 dansmith jroll: why are two things racing to the same node?
16:58:17 melwitt or, add a way to do the reserve first thing
16:58:24 dansmith jroll: scheduler should have prevented that already

Earlier   Later