| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-07-24 | |||
| 18:46:47 | dansmith | in PS10 I mean | |
| 18:47:26 | edleafe | dansmith: correct. Want me to add some safety stuff there? | |
| 18:47:57 | dansmith | edleafe: yeah, warn and skip if unset, test for those cases and then I think we can land this and argue about the batching after | |
| 18:48:38 | s-dean | mriedem: that would be brilliant. before i head home for the night, can i specify in keystone, glance and nova their database connections to be localhost skipping the need for TLS, or do they need to be specified with a network address like in the docs? do the other service such as neutron and and nova-compute need to access the DB over the network? | |
| 18:49:04 | edleafe | dansmith: ok cool. I asked in -ironic about how often the compute service would be restarted. Got one answer from jroll, and that was he only did it during upgrades | |
| 18:49:12 | edleafe | IOW, not a common thing for him | |
| 18:49:22 | mriedem | s-dean: nova-compute definitely does not use the db locally | |
| 18:49:35 | mriedem | i don't know about the various neutron agents | |
| 18:49:37 | dansmith | edleafe: heh, sure, but jroll knows what he's doing | |
| 18:49:39 | mriedem | or cinder-volume | |
| 18:49:58 | dansmith | edleafe: like I said there are other config management approaches that are pretty darn restart happy | |
| 18:50:13 | dansmith | I know of one very specifically | |
| 18:50:15 | dansmith | one that has lots of Os in it | |
| 18:53:36 | mriedem | s-dean: got it https://review.openstack.org/#/c/486724/ | |
| 18:53:43 | jaypipes | edleafe: re: separate patch, yeah, I'm thinking it's best to generate that list of alternates in the first patch (but not change the returned value from select_destinations() and then in the followup patch, change the RPC API signature and the corresponding conductor stuff | |
| 18:53:55 | s-dean | mriedem: I'm installing a general compute cloud, nothing to fancy, | |
| 18:54:03 | openstackgerrit | Merged openstack/nova master: placement: add retry tight loop claim_resources() https://review.openstack.org/486170 | |
| 18:54:06 | s-dean | mriedem: cheers :) | |
| 18:54:13 | dansmith | jaypipes: edleafe agreed | |
| 18:54:16 | mriedem | s-dean: but the nova-compute services are running on separate hosts right/ | |
| 18:54:19 | mriedem | ? | |
| 18:54:24 | s-dean | yes | |
| 18:54:29 | mriedem | nova-compute will interact with the db via nova-conductor | |
| 18:54:33 | s-dean | compute01 | |
| 18:54:38 | s-dean | ok sweet | |
| 18:54:45 | edleafe | jaypipes: ok, so the last change will go from the current return of a list of hosts to the monsterous list of hosts+alternates+allocation_candidates, right? | |
| 18:54:49 | mriedem | so if nova-conductor is on the control nodes local to the db, then you could use localhost for nova.conf for conductor | |
| 18:54:51 | mriedem | but not nova-compute | |
| 18:55:06 | mriedem | s-dean: nova-compute shouldn't even need the [database] section filled in for nova.conf | |
| 18:55:20 | s-dean | ok awsome, yeah all the scheduler conductor novncproxy etc are on 1 controller node | |
| 18:55:30 | mriedem | then localhost or 127.0.0.1 is fine for those | |
| 18:55:43 | s-dean | looks like neutron needs access to the DB from across the network :/ | |
| 18:55:59 | s-dean | ill cross that brige when i come to it i guess | |
| 18:56:02 | mriedem | i'm not familar with everything they do | |
| 18:56:08 | jaypipes | edleafe: not sure that we need to return a list of alternates instead of just a list of allocation candidates, but wondering what dansmith thinks. | |
| 18:56:12 | mriedem | s/everything/most anything/ | |
| 18:56:25 | jaypipes | edleafe: the only reason to do so would be to keep the allocation candidates entirely opaque | |
| 18:58:43 | dansmith | jaypipes: we have to keep allocation candidates | |
| 18:58:55 | dansmith | jaypipes: the cell conductor can't know what else to claim otherwise.. a hostname/uuid isn't enough | |
| 18:59:15 | jaypipes | dansmith: no, I'm referring to whether we return a list of alternate hosts PLUS the allocation candidates OR just a list of allocation candidates. | |
| 18:59:51 | dansmith | jaypipes: ah okay | |
| 19:00:02 | mriedem | would the alternate hosts be used for anything? | |
| 19:00:22 | mriedem | does conductor use them to monkey with the request spec or scheduler hints? | |
| 19:00:41 | jaypipes | mriedem: nothing other than preventing the conductor from needing to know anything about the allocation_candidates. | |
| 19:01:06 | jaypipes | mriedem: the alternate_hosts would replace the retry block in the request spec essentially. | |
| 19:01:10 | dansmith | jaypipes: if that allows conductor to be ignorant of theallocation then yes, good plan | |
| 19:01:18 | jaypipes | dansmith: k.\ | |
| 19:01:49 | jaypipes | the decision comes down to how much opacity we want those allocation_request blocks to be. :) | |
| 19:02:03 | mriedem | where is the code being proposed that needs them together? | |
| 19:02:16 | jaypipes | mriedem: edleafe's currently working on the series. | |
| 19:03:42 | edleafe | jaypipes: the way I understood it is that we would return a series of alternate hosts in case the cell had to retry a build. Each of those hosts would need its corresponding allocation_candidate so that the cell conductor could do the proper claiming. | |
| 19:04:15 | dansmith | right | |
| 19:04:24 | dansmith | however, I was thinking that: | |
| 19:04:31 | jaypipes | edleafe: yes, that's exactly correct. I was only pointing out that technically the allocation_requests contain all the information in the alternate_hosts list. | |
| 19:04:48 | dansmith | conductor would throw the first allocation candidate at placement, and then parse the result to determine which compute host it should send the rpc message to | |
| 19:05:04 | jaypipes | edleafe: but like I said, that would require the cell conductor to understand what an allocation request was (i.e. the allocation_request would no longer be opaque) | |
| 19:05:05 | dansmith | that would mean no extra list, but also opaque allocation request | |
| 19:05:13 | dansmith | jaypipes: not if ^ | |
| 19:05:33 | jaypipes | right. :) | |
| 19:05:51 | mriedem | "conductor would throw the first allocation candidate at placement" - that's the cell conductor yes? | |
| 19:05:53 | edleafe | dansmith: "parse the result"? | |
| 19:05:55 | mriedem | during a retry | |
| 19:06:18 | dansmith | edleafe: parse the result of the POST of the allocation | |
| 19:06:37 | jaypipes | I think the most appropriate return value from select_destinations() would be a list of (host, allocation_request) tuples. | |
| 19:07:07 | jaypipes | for the first item in that list, the allocation_request would be the one that had already been successfully claimed for the selected host. | |
| 19:07:21 | jaypipes | dansmith: agree? | |
| 19:07:43 | mriedem | select_destinations today returns as the first entry the one that was chosen, right? | |
| 19:07:45 | dansmith | sure that's fine, if that's how you want it to look | |
| 19:07:54 | jaypipes | mriedem: for each instance in num_instances, yes | |
| 19:08:05 | edleafe | dansmith: reportclient.claim_resources returns a boolean | |
| 19:08:11 | jaypipes | so actually, the return value needs to be list of list of that tuple. | |
| 19:08:27 | dansmith | edleafe: what's your point? | |
| 19:08:36 | jaypipes | with the outer list being for the num_instances | |
| 19:08:37 | edleafe | dansmith: what's there to parse? | |
| 19:08:56 | mriedem | a list of lists of tuples | |
| 19:08:57 | mriedem | what could go wrong | |
| 19:09:12 | edleafe | dansmith: the cell conductor would still need to "know" about the allocation_candidate structure | |
| 19:09:14 | dansmith | edleafe: the actual POST call for /allocations returns the allocation you made right? | |
| 19:09:30 | mriedem | it's not a POST | |
| 19:09:55 | edleafe | It's a PUT | |
| 19:09:58 | dansmith | edleafe: the cell conductor can, but I think it should look at the result of the http call not the thing it was passed in rpc, otherwise we've got version mess | |
| 19:10:09 | edleafe | And it returns a 204 on success | |
| 19:10:14 | dansmith | christ, whatever | |
| 19:10:15 | mriedem | yeah no content on success | |
| 19:10:32 | jaypipes | mriedem: well, tell me if you want to stop supporting num_instances > 1 and I'll gladly submit that patch ;) | |
| 19:10:40 | dansmith | okay then that clearly won't work | |
| 19:11:34 | edleafe | For each host, you get a list of (host, alloc) tuples. | |
| 19:11:50 | jaypipes | right | |
| 19:11:52 | jaypipes | ++ | |
| 19:12:00 | edleafe | On a retry in the cell, you try claiming the alloc. If that succeeds, you build on that host | |
| 19:12:09 | jaypipes | +1 | |
| 19:12:13 | dansmith | I don't love it, but I'm also not sure why we're even discussing it | |
| 19:12:17 | edleafe | If it fails, move to the next one in the list | |
| 19:12:23 | jaypipes | right, zactly.\ | |
| 19:12:55 | edleafe | jaypipes: so I don't understand why you would want to only return alloc | |
| 19:14:14 | mriedem | and just to confirm my understand, we only ever care about the list of lists for the server create case, b/c for everything else, like migrations and unshelve, it gets back the list of host states today but just takes the first one for the rpc cast to compute | |
| 19:14:18 | mriedem | *understanding | |
| 19:14:53 | jaypipes | edleafe: never mind my thought about only returning the allocations. I've been convinced that's a bad idea. | |
| 19:15:22 | edleafe | jaypipes: roger that | |
| 19:15:55 | mriedem | seems you have to have both the HostState and allocation requests because of all the filter properties and az and limits and node crap that's embedded in the HostState object | |