Earlier  
Posted Nick Remark
#openstack-nova - 2020-05-05
19:47:33 zigo So yeah, I agree with you that it'd be nice to check many things in the /healthcheck and improve it, but *after* this patch is merged.
19:48:18 zigo BTW, we can actually disable the /healthcheck, it's enabled by default.
19:48:31 zigo It'd be just by creating a file, as per disable_by_file_path = /var/run/nova/healthcheck_disable in the patch.
19:48:39 zigo Which by the way is at the wrong location, IMO...
19:48:43 zigo It should be in /etc/nova.
19:48:47 zigo I'll change the path ...
19:48:51 zigo (right away now)
19:51:41 openstackgerrit Thomas Goirand proposed openstack/nova master: Add a /healthcheck URL https://review.opendev.org/724684
19:55:21 artom gmann, zigo, FWIW I'm in favor of the /healthcheck URL (and I feel zigo's bandwidth as an operator). My concern would be around 1. documentation 2. openstack-wide uniformity
19:55:36 artom So for instance, keystone has it as well, but IIUC you can't disable there via a file
19:56:22 zigo artom: I intend to patch other projects. Neutron has already approved my patch, and got some +1 on other projects.
19:56:35 zigo I can add more /etc/<FOO>/healthcheck_disable
19:57:04 zigo As for documentation, well, maybe this goes into oslo.middleware ?
19:57:36 zigo Do you have any other place more OpenStack wide to suggest ?
19:57:40 gmann zigo: artom yeah, and there are way to add detail in response via conf, all those we can document.
19:57:50 artom zigo, well... Maybe? The admins will be reading the project guides, not oslo.middleware docs...
19:58:19 gmann zigo: but how about API is there but no DB connection? i mean DB check can be generic and can be done on oslo side
19:58:34 zigo I don't feel like we should document this on each and every project *IF* we have consistency across projects.
19:59:03 zigo gmann: That's a call for adding more code in oslo.middleware then, no ?
19:59:05 artom zigo, fair point
19:59:27 bnemec gmann: That's the sort of thing mugsie suggested a while back.
19:59:29 artom Is the docs SIG/WG still a thing? The would know...
19:59:31 gmann zigo: yeah on oslo.middleware side.
19:59:37 artom *They
19:59:50 bnemec Basically the ability to do more than just HTTP API checks, say db, messaging, maybe libvirt for Nova.
19:59:51 gmann artom: doc SIG is there.
20:00:23 artom gmann, so they'd be the ones coming up with where to put this doc, if we get openstack-wide uniformity
20:00:35 gmann bnemec: yeah, libvirt specific can be done as nova defined backend. but yeah good to have db, messaging things on oslo.middleware
20:01:05 artom gmann, but... which DB? If the deployment is multicell, just the nova_api one?
20:01:21 artom It's a rathole that I'm not sure we want to go down into, for something this simple
20:01:30 artom /healthcheck is better than nothing, even if it's not perfect
20:01:38 gmann artom: is it technical writing SIG not doc sig - https://governance.openstack.org/sigs/
20:01:48 bnemec artom: That was our conclusion too last time we discussed it.
20:02:11 bnemec If we try to boil the ocean and healthcheck all the things this will never get done at all.
20:02:51 gmann yeah deployment model is one challenge for such checks
20:03:16 artom gmann, and presumably we can improve this in the future, no?
20:03:34 artom So maybe just version is (somehow) for now, and leave is as the simple middleware thing
20:03:48 artom And eventually add /healthcheck/v2 or whatever, which can be smarter
20:04:01 artom *version it
20:04:22 gmann url can be same and configured backends by operator can enhance the checks.
20:07:50 gmann zigo: i think we need separate new config on nova side if we want to disable it right? bnemec can correct me.
20:08:09 gmann otherwise if middleware is added then it is enabled by defualt
20:08:12 gmann default
20:10:01 bnemec If we want to make healthcheck runtime configurable then it would be preferrable if we could do it in oslo. Otherwise you end up with a half dozen differen opts in different projects designed to do the same thing.
20:11:08 gmann bnemec: yeah. i am worried to have enable by default because it is unauthenticated endpoint.
20:12:21 artom gmann, what's the harm though?
20:13:08 artom Like, OK, anyone can spam that, but so can anyone spam any other API and get a flood of 403s back
20:13:33 artom The disk of DOS (if that's your concern) is greater with the second one, because keystone has to actually do work
20:13:58 artom *risk of
20:14:42 bnemec My understanding was the unauthenticated healthchecks are generally okay. Most of the stuff calling them doesn't have auth capability anyway, so you don't have much choice.
20:15:07 gmann artom: well it is still security risk to know services running on that node
20:15:28 artom gmann, security by obscurity is Real Bad (TM) ;)
20:15:49 gmann yeah i am not suggesting to make it auth but disable by default and let operator to enable it by knowing this is unauth things.
20:15:50 artom If knowing that nova-api is running somewhere is a risk, we're doing something horribly wrong
20:16:48 bnemec How would you prevent them from knowing the service is there? A simple port scan would tell them that.
20:17:07 artom gmann, off by default works for me. But that only reinforces the doc problem - admins have to know how to enable it, and more fundamentally that it's available
20:18:18 gmann artom: doc anyways we might need to tell at least it is there now in nova and refer oslo doc for all details and config options.
20:19:56 gmann and there is no harm of doing those doc of 2-3 line in all projects start implementing that and point to oslo doc for details. because we do not have any central doc place now for openstack operators/admin etc as general
20:21:36 artom gmann, fair point
20:21:58 gmann knowing service can be hacked by multiple way. but is it fine to have this unatuh endpoint for operator without no way to disable it ? may be zigo can tell.
20:22:36 gmann i felt giving the enable/disable option for them does not harm. disable by default or enable that can be disucssed whatever is more common.
20:22:47 zigo The thing is, in a typical deployment, haproxy will be the one holding the API on a public IP.
20:23:16 zigo So, an operator could disable /healthcheck on haproxy because it's in front.
20:23:37 zigo The /healthcheck from the nova-api would typically *not* be reachable from the outside.
20:25:14 gmann but it can be reachable to anyone for other deployment which are not under haproxy or other layer of direction/load balancer etc
20:25:39 gmann disabling on haproxy side is all good i think.
20:26:04 zigo We have the / which needs to be enabled unauth anyways...
20:27:34 artom gmann, I still strongly disagree that it's a legit security risk - or if it is, we've messed up badly
20:28:01 artom Our security should be that nova-api is safe, not that we need to hide it :)
20:31:11 dansmith I haven't been paying attention here
20:31:24 dansmith but anything that generates db or mq traffic should be authenticated, I think
20:31:30 dansmith even if read-only
20:31:36 artom dansmith, it doesn't
20:32:03 dansmith okay, so a only a healthcheck of "can you respond to http requests" then?
20:32:12 artom dansmith, yep
20:32:18 dansmith because generally those expand to "can you talk to your database?" etc
20:32:49 artom dansmith, yeah, gmann wanted to do that, but IMO it's complexity we don't necessarily need (right now)
20:33:02 artom For instance, which database? api? cell1? What about multicell?
20:33:36 dansmith artom: right, so if it's going to grow to that (which most people will want it it, AFAIK) you should consider it needing to be auth'd
20:34:04 gmann dansmith: yeah, my idea is to expand that to DB checks also
20:34:06 artom dansmith, no argument there from me
20:34:07 dansmith because I would expect returning a dict of each db (api, cell1, cell2, etc) as well as each mq would be exactly what people want
20:34:16 dansmith gmann: right
20:34:38 dansmith artom: so are you arguing that if you have no auth, you get a "yes I'm here but that's all I will tell you" response/
20:36:04 gmann artom: and i will say we finalize it now what all things we want to do with healthcheck instead of expending it later which make lot of things change like auth, enable/disable things etc. but only https request checks we are like 'everything is fine as it is'
20:36:10 artom dansmith, it was more of a response to gmann's security risk argument - that just knowing a service is running somewhere is not a security risk, or if it is, we've done something horribly wrong
20:36:48 artom gmann, you mean future proofing it?
20:37:08 bnemec We've been discussing them for years and nobody has actually implemented anything.
20:37:16 gmann that is why i suggested to have spec and then discuss what all things we want to do with that.
20:37:49 dansmith gmann: we should never expect to design something like that once and forever, but agree we should not design it before making it useful, only to have to revise
20:37:55 artom gmann, yeah, I get that... It's just... looks to me like it'll get bogged down in the kind of discussion that we're having :)
20:38:27 gmann bnemec: yeah, it has more attention now as zigo also brought it up for doing it for all projects as community goal which is good idea.
20:38:42 artom So I'd be more in the campg of "do something easy and useful *now*, trying to future-proof it", instead "do something perfect later"
20:38:44 bnemec It's not the first time that has happened either. I proposed it in Berlin as a community goal.
20:38:46 dansmith artom: so you leave all your computer's ports open while connected directly to the internet and depend on windows/samba auth to keep people out of your files right? :)
20:39:23 dansmith expose ssh to everyone, relying on your password and the patched-ness of your sshd, etc
20:40:12 artom I mean, I've almost done the latter
20:40:34 gmann artom: but it can break things later say, it is enable by default and if we extend in future for things need more care then we may want to disable it by default.
20:40:54 gmann bnemec: ah, did not know that :)
20:40:59 artom gmann, I'm not arguing off by default :)

Earlier   Later