Earlier  
Posted Nick Remark
#openstack-nova - 2019-10-30
20:52:21 dansmith efried: asking because I think I've seen some notes from you around this code in the past,
20:52:41 dansmith but do you have strong opinions on what we should do if we have, say, three glance endpoints and we get an error from one trying to do something?
20:52:57 dansmith looks like we round-robin the available endpoints, but don't move onto the next one after a failure
20:52:57 efried dansmith: we should not have three glance endpoints
20:53:09 efried unless you mean three interfaces
20:53:20 dansmith efried: three urls
20:53:23 efried f, did we not get rid of that api_servers thing yet?
20:53:25 efried stand by
20:53:35 dansmith no, the round-robin-ing you mean? that's still there
20:54:10 dansmith I honestly thought that we had pushed responsibility for that down into the glance client, but ... the round-robin stuff is there
20:54:12 efried omg we haven't even deprecated it yet!
20:54:22 efried we should do that.
20:54:37 dansmith efried: use more words.. what are the people currently using that supposed to do?
20:54:55 efried sorry
20:54:56 efried okay
20:55:27 efried people using api_servers were theoretically stuffing multiple endpoints in there, and then they were responsible for making sure all of those endpoints pointed to the same images.
20:55:37 dansmith right
20:55:41 efried nova would, on a given call, pick "the next one" and use it.
20:56:00 efried afaik nova never had any logic to, say, "try the next one" if one failed.
20:56:13 dansmith right, it doesn't, but if we can have multiples we should, but.. go on
20:56:18 efried Because nova doesn't keep track of how many there are. It just loads them up and cycles over them.
20:56:40 efried so if we had only one, we still "cycle" over just the one.
20:56:50 efried and if we tried to add a "try next" thing, we would just end up trying... the same one.
20:57:28 dansmith well, all we'd have to do is retry N-1 times for an endpoint count of N and then we'd hit them all and not dupe one or more
20:57:29 efried if api_servers was a thing we still wanted people to use, I could see enhancing our logic to do that properly - keep track of how many there are and, on certain classes of error, try the next and so on until we succeed or come full circle
20:57:31 efried but
20:57:39 efried we don't want people to use api_servers.
20:57:47 dansmith and that is because why?
20:57:49 efried They should use a single endpoint that's a front for a load balancer
20:58:02 efried and use standard ksa options like every other service
20:58:25 dansmith yeah, so the problem there is that you have to have a HA LB for that
20:59:04 dansmith cloud client software needs to handle failures like this to be robust, I'm not sure why nova shouldn't
20:59:11 efried I thought there was a project that did that for you
20:59:26 efried do we try to HA other things? cinder, keystone, neutron...?
20:59:33 efried (trick question)
20:59:52 dansmith well, glance and keystone would be much easier to retry than the others
20:59:59 smcginnis Not the API, but cinder services can be run active/active.
21:00:22 smcginnis And I've seen people use the active/passive pacemaker setup for things too.
21:00:29 dansmith yeah
21:00:31 efried smcginnis: behind a single catalog endpoint tho?
21:00:36 smcginnis Yeah
21:00:39 efried melwitt: would you mind giving this another look? https://review.opendev.org/#/c/615704/
21:00:52 dansmith efried: well, point is, if you're going to deprecate it, I'd appreciate you getting that patch up so I can point the people I'm talking to about it
21:01:13 dansmith efried: I remember you talking about this a while ago, had kinda assumed it had already happened
21:01:32 efried dansmith: are you playing devil's advocate or do you really think nova should be responsible for load balancing glance?
21:01:34 dansmith tripleo is using api_servers, so they're going to need some notice
21:02:05 dansmith efried: not load balancing, but being able to have multiple endpoints and find one that works, yeah
21:02:33 efried yeah, "everyone" is probably still using api_servers, because we only added the standard ksa stuff in, I think, queens.
21:02:41 dansmith having to have a floating VIP and HA'ing services behind that VIP for everything really sucks
21:02:59 efried is tripleo using multiple endpoints to api_servers?
21:03:12 dansmith efried: just to be clear, "the ksa stuff" doesn't imply anything about multiple endpoints right?
21:03:15 dansmith efried: they are
21:03:19 melwitt efried: whoa, that is old. I can look at it again once I figure out what it's about again
21:03:37 dansmith efried: for glance, but we don't do the failover stuff, so not really for much gain
21:03:38 efried melwitt: thanks, I got a nudge from a downstream (not $employer) who needs it.
21:03:47 melwitt ack
21:04:02 efried dansmith: right, the way nova handles it really isn't HA-ish, it's more load-balance-ish.
21:04:29 dansmith efried: currently, right
21:04:51 efried dansmith: and no, "the ksa stuff" doesn't have anything for mulitiple endpoints... I don't think. (Can you put multiple endpoints in the service catalog?)
21:05:25 dansmith gah
21:05:30 efried dansmith: I can tell you this, cutting over to sdk to talk to glance, api_servers will have to be gone first.
21:05:46 dansmith brb
21:06:34 dansmith efried: I dunno about the service catalog, probably not
21:06:43 dansmith not saying that's a good thing, but.. :)
21:07:10 dansmith efried: okay so, can you put up a patch to deprecate it real quick so I can point people to it?
21:07:17 efried I'm going to have to nail down mordred for the more philosophical aspects of this.
21:07:20 dansmith efried: -W it if you want for completeness and comments or something
21:07:37 efried I can do that, yeah, though now you've got me doubting if we *should*.
21:07:38 dansmith efried: ack, we can blame lots of stuff on him like this
21:07:47 dansmith efried: that's fine, a place to collect feedback
21:07:50 efried ...
21:08:09 dansmith efried: the tripleo people can go voice their hate there if they want, or they may say "meh, okay, we can VIP this too" I honestly don't know
21:08:20 efried cool
21:08:39 dansmith efried: can you do that, like, uh, quickly so I can reply to someone quickly? :D
21:08:47 efried on it rn
21:08:52 dansmith spanks.
21:08:56 efried if people would stop bugging me in IRC.
21:12:57 efried dansmith: while you're waiting, I found this (search for api_servers) http://specs.openstack.org/openstack/nova-specs/specs/queens/implemented/use-ksa-adapter-for-endpoints.html#proposed-change
21:13:47 dansmith "The exception is [glance]api_servers, which will continue to be supported."
21:14:51 dansmith I wonder if "valid_interfaces" is an arbitrary list, such that you could have internal1, internal2, internal3?
21:15:48 openstackgerrit Joshua Cornutt proposed openstack/nova master: Moving get_hash_str() from md5 to sha-256 https://review.opendev.org/615704
21:16:20 dansmith I also don't know how LBs (or HAproxy) handle a failed http session post-request.. meaning if I do a GET against glance and then the connection is reset after I've sent the GET, does the LB retry and re-play that against the next server?
21:16:24 mriedem dansmith: i think it might be, we've had to dig into the ksa code a few times to answer that
21:16:44 dansmith I'd kindof doubt it.. figure it'd fail the client, mark the server, next request doesn't include the bad server
21:16:56 dansmith which means if we don't retry we propagate the failure up
21:17:18 dansmith which is fine in the trivial case, but not so much when we're making multiple calls as part of a complex operation, like against cinder or neutron
21:17:39 dansmith mriedem: might be what, arbitrary?
21:17:45 dansmith or might be constrained?
21:18:32 dansmith based on the description in that spec, though, it tries them in order, so you'd hammer internal1 until it fails, and then try internal2, it sounds like
21:18:37 dansmith which wouldn't be what you'd want really
21:21:26 openstackgerrit Eric Fried proposed openstack/nova master: Deprecate [glance]api_servers https://review.opendev.org/692227
21:21:28 mriedem dansmith: i think valid_interfaces is arbitrary but i've had to look this up multiple times in the ksa code
21:21:28 efried dansmith: ^
21:21:38 dansmith efried: thanks
21:21:50 efried dansmith: no, valid_interfaces is public/private/internal only
21:21:59 efried I think they can have 'URL' suffixes for legacy reasons
21:22:01 mriedem also, on that note in the spec for api_servers, i want to say the godaddy or smorrison made a point that they still really needed that when the spec was being reviewed
21:22:48 dansmith mriedem: yeah, I *think* people are going to be quite unhappy about it, but we'll see
21:22:56 mriedem might have been in the ML, i don't see it in review comments
21:22:59 efried "really needed" meaning it was easier for us to continue to support it than it was for them to stop relying on our shoddy load balancing hack

Earlier   Later