Earlier  
Posted Nick Remark
#openstack-nova - 2017-07-24
18:47:57 dansmith edleafe: yeah, warn and skip if unset, test for those cases and then I think we can land this and argue about the batching after
18:48:38 s-dean mriedem: that would be brilliant. before i head home for the night, can i specify in keystone, glance and nova their database connections to be localhost skipping the need for TLS, or do they need to be specified with a network address like in the docs? do the other service such as neutron and and nova-compute need to access the DB over the network?
18:49:04 edleafe dansmith: ok cool. I asked in -ironic about how often the compute service would be restarted. Got one answer from jroll, and that was he only did it during upgrades
18:49:12 edleafe IOW, not a common thing for him
18:49:22 mriedem s-dean: nova-compute definitely does not use the db locally
18:49:35 mriedem i don't know about the various neutron agents
18:49:37 dansmith edleafe: heh, sure, but jroll knows what he's doing
18:49:39 mriedem or cinder-volume
18:49:58 dansmith edleafe: like I said there are other config management approaches that are pretty darn restart happy
18:50:13 dansmith I know of one very specifically
18:50:15 dansmith one that has lots of Os in it
18:53:36 mriedem s-dean: got it https://review.openstack.org/#/c/486724/
18:53:43 jaypipes edleafe: re: separate patch, yeah, I'm thinking it's best to generate that list of alternates in the first patch (but not change the returned value from select_destinations() and then in the followup patch, change the RPC API signature and the corresponding conductor stuff
18:53:55 s-dean mriedem: I'm installing a general compute cloud, nothing to fancy,
18:54:03 openstackgerrit Merged openstack/nova master: placement: add retry tight loop claim_resources() https://review.openstack.org/486170
18:54:06 s-dean mriedem: cheers :)
18:54:13 dansmith jaypipes: edleafe agreed
18:54:16 mriedem s-dean: but the nova-compute services are running on separate hosts right/
18:54:19 mriedem ?
18:54:24 s-dean yes
18:54:29 mriedem nova-compute will interact with the db via nova-conductor
18:54:33 s-dean compute01
18:54:38 s-dean ok sweet
18:54:45 edleafe jaypipes: ok, so the last change will go from the current return of a list of hosts to the monsterous list of hosts+alternates+allocation_candidates, right?
18:54:49 mriedem so if nova-conductor is on the control nodes local to the db, then you could use localhost for nova.conf for conductor
18:54:51 mriedem but not nova-compute
18:55:06 mriedem s-dean: nova-compute shouldn't even need the [database] section filled in for nova.conf
18:55:20 s-dean ok awsome, yeah all the scheduler conductor novncproxy etc are on 1 controller node
18:55:30 mriedem then localhost or 127.0.0.1 is fine for those
18:55:43 s-dean looks like neutron needs access to the DB from across the network :/
18:55:59 s-dean ill cross that brige when i come to it i guess
18:56:02 mriedem i'm not familar with everything they do
18:56:08 jaypipes edleafe: not sure that we need to return a list of alternates instead of just a list of allocation candidates, but wondering what dansmith thinks.
18:56:12 mriedem s/everything/most anything/
18:56:25 jaypipes edleafe: the only reason to do so would be to keep the allocation candidates entirely opaque
18:58:43 dansmith jaypipes: we have to keep allocation candidates
18:58:55 dansmith jaypipes: the cell conductor can't know what else to claim otherwise.. a hostname/uuid isn't enough
18:59:15 jaypipes dansmith: no, I'm referring to whether we return a list of alternate hosts PLUS the allocation candidates OR just a list of allocation candidates.
18:59:51 dansmith jaypipes: ah okay
19:00:02 mriedem would the alternate hosts be used for anything?
19:00:22 mriedem does conductor use them to monkey with the request spec or scheduler hints?
19:00:41 jaypipes mriedem: nothing other than preventing the conductor from needing to know anything about the allocation_candidates.
19:01:06 jaypipes mriedem: the alternate_hosts would replace the retry block in the request spec essentially.
19:01:10 dansmith jaypipes: if that allows conductor to be ignorant of theallocation then yes, good plan
19:01:18 jaypipes dansmith: k.\
19:01:49 jaypipes the decision comes down to how much opacity we want those allocation_request blocks to be. :)
19:02:03 mriedem where is the code being proposed that needs them together?
19:02:16 jaypipes mriedem: edleafe's currently working on the series.
19:03:42 edleafe jaypipes: the way I understood it is that we would return a series of alternate hosts in case the cell had to retry a build. Each of those hosts would need its corresponding allocation_candidate so that the cell conductor could do the proper claiming.
19:04:15 dansmith right
19:04:24 dansmith however, I was thinking that:
19:04:31 jaypipes edleafe: yes, that's exactly correct. I was only pointing out that technically the allocation_requests contain all the information in the alternate_hosts list.
19:04:48 dansmith conductor would throw the first allocation candidate at placement, and then parse the result to determine which compute host it should send the rpc message to
19:05:04 jaypipes edleafe: but like I said, that would require the cell conductor to understand what an allocation request was (i.e. the allocation_request would no longer be opaque)
19:05:05 dansmith that would mean no extra list, but also opaque allocation request
19:05:13 dansmith jaypipes: not if ^
19:05:33 jaypipes right. :)
19:05:51 mriedem "conductor would throw the first allocation candidate at placement" - that's the cell conductor yes?
19:05:53 edleafe dansmith: "parse the result"?
19:05:55 mriedem during a retry
19:06:18 dansmith edleafe: parse the result of the POST of the allocation
19:06:37 jaypipes I think the most appropriate return value from select_destinations() would be a list of (host, allocation_request) tuples.
19:07:07 jaypipes for the first item in that list, the allocation_request would be the one that had already been successfully claimed for the selected host.
19:07:21 jaypipes dansmith: agree?
19:07:43 mriedem select_destinations today returns as the first entry the one that was chosen, right?
19:07:45 dansmith sure that's fine, if that's how you want it to look
19:07:54 jaypipes mriedem: for each instance in num_instances, yes
19:08:05 edleafe dansmith: reportclient.claim_resources returns a boolean
19:08:11 jaypipes so actually, the return value needs to be list of list of that tuple.
19:08:27 dansmith edleafe: what's your point?
19:08:36 jaypipes with the outer list being for the num_instances
19:08:37 edleafe dansmith: what's there to parse?
19:08:56 mriedem a list of lists of tuples
19:08:57 mriedem what could go wrong
19:09:12 edleafe dansmith: the cell conductor would still need to "know" about the allocation_candidate structure
19:09:14 dansmith edleafe: the actual POST call for /allocations returns the allocation you made right?
19:09:30 mriedem it's not a POST
19:09:55 edleafe It's a PUT
19:09:58 dansmith edleafe: the cell conductor can, but I think it should look at the result of the http call not the thing it was passed in rpc, otherwise we've got version mess
19:10:09 edleafe And it returns a 204 on success
19:10:14 dansmith christ, whatever
19:10:15 mriedem yeah no content on success
19:10:32 jaypipes mriedem: well, tell me if you want to stop supporting num_instances > 1 and I'll gladly submit that patch ;)
19:10:40 dansmith okay then that clearly won't work
19:11:34 edleafe For each host, you get a list of (host, alloc) tuples.
19:11:50 jaypipes right
19:11:52 jaypipes ++
19:12:00 edleafe On a retry in the cell, you try claiming the alloc. If that succeeds, you build on that host
19:12:09 jaypipes +1
19:12:13 dansmith I don't love it, but I'm also not sure why we're even discussing it
19:12:17 edleafe If it fails, move to the next one in the list
19:12:23 jaypipes right, zactly.\
19:12:55 edleafe jaypipes: so I don't understand why you would want to only return alloc
19:14:14 mriedem and just to confirm my understand, we only ever care about the list of lists for the server create case, b/c for everything else, like migrations and unshelve, it gets back the list of host states today but just takes the first one for the rpc cast to compute
19:14:18 mriedem *understanding
19:14:53 jaypipes edleafe: never mind my thought about only returning the allocations. I've been convinced that's a bad idea.
19:15:22 edleafe jaypipes: roger that
19:15:55 mriedem seems you have to have both the HostState and allocation requests because of all the filter properties and az and limits and node crap that's embedded in the HostState object
19:16:00 mriedem which conductor is using before casting to compute
19:16:02 mriedem yeah?

Earlier   Later