| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-04-24 | |||
| 16:40:25 | melwitt | I mean, the cheat would be move that call to "prepare networking etc" method | |
| 16:40:27 | TheJulia | https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L360 | |
| 16:40:59 | jroll | melwitt: ya, that's the other version: https://review.openstack.org/#/c/563714/ | |
| 16:41:05 | TheJulia | :) | |
| 16:41:07 | jaypipes | dansmith: no worries. we were discussing that ironic (actually, some cinder drivers) need to know the IP address of a port up front. and the prep_network_for-blokc_devices() virt driver API method was added in order to address that need. But that method does vif plugging (because that's unfortunately when a port's IP address is guaranteed to be set). We would like to see the IP address allocation decoupled from the vif plugging action. | |
| 16:41:22 | jroll | feels out of place to me but if the bug gets fixed then ¯\_(ツ)_/¯ | |
| 16:41:26 | melwitt | jroll: ah, so that's what y'all mean by "lock" | |
| 16:41:44 | dansmith | jaypipes: okay | |
| 16:41:46 | TheJulia | s/lock/giant flag saying node is in use/ | |
| 16:42:00 | jroll | melwitt: we also have internal locks which block certain actions on the node, so... sometimes :) | |
| 16:42:10 | melwitt | yeah, makes sense | |
| 16:42:18 | melwitt | heh | |
| 16:42:50 | melwitt | jaypipes, dansmith: but it can't be because DHCP, right? | |
| 16:43:01 | dansmith | melwitt: hmm? | |
| 16:43:26 | dansmith | melwitt: once you know the host, you can do the binding and get a port allocation at that point | |
| 16:43:36 | melwitt | dansmith: like, I had been thinking up to now that we get an IP allocated when we create the neutron port. but in the case of DHCP being used, that would not be true, right? | |
| 16:43:37 | dansmith | er, an address allocation for the port I mean | |
| 16:43:42 | jaypipes | melwitt: yes, when a port is in a subnet that is DHCP-enabled, the port does not get an IP address until after vif plugging (and the DHCP lease is done) | |
| 16:44:16 | dansmith | jaypipes: I think you get an address when you host-bind it, not exactly plug right? | |
| 16:44:23 | jroll | note that ironic doesn't bind the host to the port in nova, but later in ironic, because it can't be attached to the tenant network until it's deployed | |
| 16:45:01 | melwitt | dansmith: because ultimately what they need is to know the IP before they attach the volume. and so far, they have to plug the vif first in order to get the IP. and the problem that's happening is that the vif plug is happening outside of the node reserve lock, so things are racing | |
| 16:45:02 | jaypipes | dansmith: in the case of DHCP, the VIF needs to be fully set up and then the DHCP request made to the gateway, though, right? only after that will the port get an IP adderess. | |
| 16:45:28 | dansmith | jaypipes: you don't have to hit the dhcp server with a client to get an address | |
| 16:46:01 | jaypipes | dansmith: sorry, I wasn't aware that was possible. | |
| 16:46:08 | dansmith | melwitt: I'm not positive at which exact step neutron will assign an address I guess (plug vs. bind), | |
| 16:46:33 | dansmith | but I'm nearly positive it happens before the client is really set up, | |
| 16:46:47 | dansmith | otherwise you wouldn't be able to see what ip your instance is going to have until it has come up enough to have hit the dhcp server | |
| 16:46:47 | melwitt | dansmith: plug is just a local thing (os-vif in the case of libvirt) so it must be the bind, I think | |
| 16:46:50 | dansmith | which wouldn't make any sense | |
| 16:47:07 | melwitt | sean-k-mooney we need you! | |
| 16:47:10 | melwitt | :) | |
| 16:47:27 | jroll | I feel like IP allocation is done at port create time, but I would need to verify | |
| 16:47:35 | TheJulia | I'm 95% sure it is | |
| 16:47:38 | dansmith | jroll: exactly | |
| 16:47:43 | melwitt | that's what I said earlier | |
| 16:47:54 | jroll | dansmith: with or without a host binding, to be clear | |
| 16:47:57 | melwitt | so I was thinking, why not just ask neutron for the IP instead of doing plug_vifs | |
| 16:48:05 | melwitt | *plug_vifs early | |
| 16:48:15 | dansmith | jroll: I think it depends on how the network is setup whether you get it early or late, IIRC, but I dunno | |
| 16:48:25 | jaypipes | dansmith: by "port creation time" are you referring to when neutron port-create is done? because I'm pretty sure that *isn't* when IP allocation is done. | |
| 16:48:33 | jroll | could be, neutron is just a framework after all :) | |
| 16:48:49 | dansmith | I think it can happen at port-create time, and I think it can happen at host bind time | |
| 16:48:56 | dansmith | but I don't think it happens at vif_plug time | |
| 16:48:57 | melwitt | yes it's configurable https://developer.openstack.org/api-ref/network/v2/#ip-allocation-extension | |
| 16:49:27 | dansmith | think about when you're booting an instance and when you get to see the IP via the api in the scheme of it booting | |
| 16:49:29 | dansmith | usually you see it before it has even finished downloaded the image to the compute node right? | |
| 16:49:40 | melwitt | right, that's what I thought | |
| 16:49:58 | melwitt | but I thought is that only for static IPs and not DHCP? | |
| 16:50:04 | melwitt | I had thought it didn't matter | |
| 16:50:15 | dansmith | dhcp is just how you communicate the ip to the guest, | |
| 16:50:26 | dansmith | I don't think it changes how/when the port would get assigned an ip | |
| 16:50:51 | jroll | melwitt: so, ironic does a very late host binding of the port, because at that time it's put onto the tenant network (which we don't want during deployment). our plug_vifs call is the api endpoint for the code that does the host binding in ironic: https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L1620 | |
| 16:51:35 | melwitt | yeah, good point. we use DHCP in the gate and yeah, pretty sure the IP is known once the neutron port is created. but we do know that's configurable, not necessarily true that the IP will be known at port create time depending on how the network is setup in neutron, seems like | |
| 16:52:08 | jroll | (and so during BFV deployment, prepare_networks_before_block_device_mapping is a fine time to put it on the tenant network, since we skip the deployment ramdisk) | |
| 16:52:18 | melwitt | jroll: yeah, figured it must be because plug_vifs was needed to get the IP | |
| 16:53:04 | jroll | melwitt: I assume it's just to hook up the networks, but not sure, I don't know this code well | |
| 16:53:09 | artom | Which one between flavor extra specs and image properties is arbitrary again? | |
| 16:53:17 | dansmith | artom: the former | |
| 16:53:30 | dansmith | melwitt: maybe mlavalle is around and could answer some questions | |
| 16:53:57 | artom | dansmith, so operators can add a "foo" extra spec, and enable a filter (which one?) that would schedule all those foos to a certain aggregate? | |
| 16:53:57 | jroll | ah yes, the IP comes from the BDM info: https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L399 | |
| 16:54:12 | jroll | the plug_vifs call is just to get networking up | |
| 16:54:19 | melwitt | jroll: yeah, so back when we started talking about this today I was saying can we decouple the vif plugging from the IP get query and leave the vif plugging for after the node reserve like it used to be. not sure if the host bind should be behind reserve too though | |
| 16:54:54 | dansmith | artom: I have to go look for the linkage every time I'm asked.. I think there is an AggregateExtraSpecs filter or something that you use | |
| 16:54:57 | jroll | melwitt: the vif plugging is the host bind for us, but yeah, good question | |
| 16:55:16 | artom | dansmith, aha, thanks, I'll dig in the code then | |
| 16:55:19 | melwitt | jroll: this is the change that moved vif plug out from the node reserve in order to get an IP for the volume connector https://review.openstack.org/#/c/468353/19/nova/virt/ironic/driver.py | |
| 16:55:20 | dansmith | see I think in the late case, the host binding step is when you get your allocation | |
| 16:55:24 | jroll | melwitt: it feels like we could | |
| 16:56:06 | melwitt | dansmith: yeah. so would it be safe to do that outside of node reserve? jroll? | |
| 16:56:26 | jroll | dansmith: melwitt: oh, right, that docstring tells us exactly that | |
| 16:56:26 | melwitt | besides that, wouldn't that require a change to ironic API too? | |
| 16:56:54 | dansmith | melwitt: you can do host binding before spawn, that should be fine, you just can't do the plug before it | |
| 16:57:15 | melwitt | k | |
| 16:57:31 | jroll | melwitt: so, the problem we're seeing is when scheduling races with two instances to a node, they both try to do the plug because we haven't set that reservation. so I'm thinking if we do that reservation first thing, then we never hit this again | |
| 16:58:02 | melwitt | so it sounds like we have two options: decouple the host binding and do that in prepare_networks_before_block_device_mapping, get the IP for the volume attach, then reserve, then plug vifs etc | |
| 16:58:17 | melwitt | or, add a way to do the reserve first thing | |
| 16:58:17 | dansmith | jroll: why are two things racing to the same node? | |
| 16:58:24 | dansmith | jroll: scheduler should have prevented that already | |
| 16:58:37 | melwitt | dansmith: the reserve isn't done until spawn | |
| 16:58:44 | dansmith | melwitt: doesn't matter | |
| 16:58:47 | jroll | dansmith: not a clue | |
| 16:58:51 | melwitt | hm | |
| 16:58:55 | TheJulia | and there can be multiple consumers of ironic's api | |
| 16:59:01 | dansmith | melwitt: we have committed the allocation in scheduler, which means the node isn't a candidate for any other scheduler runs | |
| 16:59:13 | TheJulia | so we can't rely upon nova's scheduler to be the source of truth on locking | |
| 16:59:18 | dansmith | TheJulia: is that it? nova is racing with something else? | |
| 16:59:37 | melwitt | but is that what's happening in the environment you're seeing the problem? multiple consumers of ironic API? | |
| 16:59:44 | TheJulia | dansmith: nova can race with something else, nova can also race if resource classes are not used, which is how we first noticed the issue actually | |
| 16:59:47 | jroll | oh my | |
| 17:00:15 | dansmith | well, "if resource classes are not used" should not be a thing | |
| 17:00:24 | TheJulia | agreed, totally agreed | |
| 17:00:27 | dansmith | but if it's the former, then that's fine | |
| 17:00:39 | jroll | dansmith: I wasn't aware the allocation in scheduler landed, my bad | |
| 17:01:10 | dansmith | jroll: yeah, so nova racing with something else is a thing, I guess, but makes it less concerning to me | |
| 17:02:06 | melwitt | yeah, I mean really I feel like we could just change this to something more generic, "pre-spawn-prep" kind of method and use that to do a node reserve https://review.openstack.org/#/c/468353/19/nova/virt/driver.py | |
| 17:02:20 | TheJulia | jroll: there are some discussions in the hpc/htc world of central ironic, multiple consumers | |
| 17:02:27 | jroll | dansmith: but anyway, I thought there was still an opportunity for a scheduler race, glad to here that's gone | |
| 17:02:31 | dansmith | melwitt: gotta make sure you have a post-spawn-failure method to unreserve too | |
| 17:02:38 | TheJulia | like, ironic ends up being a central broker of baremetal across datacenters *blink* *blink* | |