Earlier  
Posted Nick Remark
#openstack-nova - 2017-07-24
19:23:58 mriedem we agree we need allocation requests and the host_state dict thingies i think
19:24:09 mriedem or maybe we don't agree
19:24:13 dansmith I think he's saying we need to pass the HostState mess, right mriedem ?
19:24:16 mriedem yes
19:24:25 jaypipes is even used....
19:24:28 mriedem it is
19:24:29 mriedem in the claim code
19:24:32 dansmith oh it is
19:24:33 dansmith yeah
19:24:43 mriedem note that all of the docstrings say the limits are vcpus/ram/disk
19:24:45 mriedem not numa
19:24:48 mriedem so it's totally f'ing confusing
19:25:07 mriedem no desert eagle?
19:25:14 mriedem you want open casket?
19:25:15 jaypipes yeah, it is. :(
19:25:23 dansmith mriedem: it's what I have within reach
19:25:53 jaypipes dansmith, mriedem: k, so the returned value needs to be list of lists of (host_dict_with_limits_thing, alloc_request)
19:25:57 jaypipes edleafe: ^
19:26:13 dansmith no
19:26:14 dansmith because
19:26:26 dansmith host_dict_with_limits is an unversioned structure and we're not adding a parameter with a LIST OF TUPLES OF THAT THING to one of our clean rpc calls
19:26:51 mriedem clean rpc calls?
19:27:03 jaypipes dansmith: don't we already pass the limits stuff down to build_instance()?
19:27:13 mriedem select_destinations today returns the unversioned host_dict_with_limits_thing
19:27:17 dansmith jaypipes: no we get it fromthe scheduler
19:27:25 mriedem jaypipes: the NUMATopologyFilter
19:27:52 jaypipes mriedem: no, I know that. I meant where does the limits get passed to instance_claim() on the compute host.
19:27:54 mriedem jaypipes: let me present exhibit Z98 https://review.openstack.org/#/c/484439/
19:28:09 mriedem via the filter props
19:28:21 mriedem actually no,
19:28:25 mriedem build_and_run_instance has a limits kwarg
19:28:33 mriedem which is the limits from the scheduler
19:28:46 mriedem which as far as i know just contains the numa topology limits
19:28:47 mriedem EXCEPT
19:28:54 mriedem for out of tree scheduler drivers that stash crap in there too :)
19:28:58 mriedem *and filters
19:29:06 dansmith hmm
19:29:32 dansmith so maybe we talk to scheduler late enough in this process that we're only calling ugly rpc calls with limits as a param from here on out?
19:29:48 mriedem do we need a hangout?
19:29:55 dansmith I need a therapist.
19:30:07 mriedem laura's mom went to grad school for that i think, she'll talk your ear off
19:30:21 mriedem about vacuums that also polish wood floors
19:30:26 mriedem controlled via your iphone app
19:30:38 jaypipes I'm wondering if we can't just recreate the limits['numa_topology'] on the compute host....
19:31:16 mriedem i guess i'm failing to see the issue with sending both back from select_destinations
19:31:25 mriedem we're already sending host_dict_with_limits_thing
19:31:35 jaypipes dansmith: ^
19:31:37 mriedem we're just tacking allocation_request(s)? onto that
19:31:39 mriedem as a tuple
19:31:55 mriedem btw, are these allocation requests plural or singular?
19:32:05 edleafe mriedem: we're changing 2 things
19:32:05 mriedem one per host right? so singular
19:32:10 jaypipes mriedem: singular per host, yeah.
19:32:12 edleafe instead of a single host
19:32:13 dansmith yeah, like I said above, I think I was missing that we're already passing that grossness in the rpc calls downstream from where we get them
19:32:21 edleafe we're sending a list of hosts
19:32:34 edleafe and each of those has an associated allocation_candidate
19:32:35 mriedem edleafe: please provide context
19:32:42 mriedem "sending" from where to where?
19:32:51 edleafe to the cell conductor
19:32:57 edleafe from the super conductor
19:33:10 dansmith super conductor doesn't call cell conductor
19:33:19 dansmith super conductor calls the first compute node, which will call cell conductor on retry
19:33:35 mriedem edleafe: ok so here https://github.com/openstack/nova/blob/master/nova/conductor/manager.py#L1046
19:33:44 edleafe I thought that was changing so that we supported alternates
19:34:17 dansmith no
19:34:35 edleafe So we're sending the compute node the big honking list of list of stuff?
19:34:42 mriedem api -> super conductor > scheduler > super conductor > compute > cell conductor (retry) > compute
19:34:52 dansmith right
19:34:54 edleafe And then it sends it to cell conductor on retry?
19:34:59 mriedem yes
19:35:05 mriedem the allocation requests are getting passed through
19:35:13 mriedem like barnacles
19:35:23 mriedem or a tube sock in mortier
19:35:26 mriedem *mortimer
19:35:27 edleafe more like kidney stones
19:35:45 mriedem 2nd question,
19:35:48 melwitt lol omg, mortimer
19:36:12 mriedem are we changing the build_and_run_instances compute rpc api to pass allocation requests, or shoving those into something else already being sent, like request spec or filter properties?
19:37:37 dansmith one of those
19:37:58 mriedem which would also impact the build_and_run_instance method in conductor rpc api
19:38:11 mriedem *build_instances
19:38:35 dansmith if we put it into something like reqspec, we won't be able to send them to older computes
19:38:41 dansmith which will break our upgrade process
19:38:46 dansmith because they'll kick the newer version back
19:39:00 mriedem personally i think it's cleaner as a new parameter on the rpc api
19:39:08 dansmith which I guess is the same for the new param approach
19:39:26 dansmith we just need to handle the case in the retry logic,
19:39:35 dansmith if we didn't get these new things, assume we can talk to the scheduler and do a reschedule
19:40:18 mriedem but if your cell conductor is blocked from up calls to the scheduler, how would that work?
19:40:31 dansmith if you have old computes,
19:40:39 dansmith then you don't have a multi-tier cellsv2 environment,
19:40:43 dansmith because we didn't support it before,
19:40:46 openstackgerrit Ed Leafe proposed openstack/nova master: Migrate Ironic Flavors https://review.openstack.org/484949
19:40:46 dansmith thus it must be fine
19:41:11 mriedem idk, we don't really have anything doc'ed for this
19:41:21 dansmith eh?
19:41:24 dansmith it can't work right now
19:41:29 dansmith no need to doc it :)
19:41:40 mriedem what does old computes have to do with multi-tier cells v2?

Earlier   Later