| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-14 | |||
| 13:34:51 | openstackgerrit | Rodolfo Alonso Hernandez proposed openstack/os-vif master: Migration from ``ip`` commands to ``pyroute2`` https://review.openstack.org/484386 | |
| 14:29:52 | pooja | Hi.. I had a question around configuring nova-scheduler in the control plane. For scalability, is it possible to run multiple nova-scheduler processes like api and conductor? | |
| 14:30:51 | pooja | As I understand, the host manager needs to be aware of all provisioning ops so it may not be possible. Any thoughs? Thanks! | |
| 14:54:43 | johnthetubaguy | gmann: here we go https://review.openstack.org/#/c/435484 | |
| 14:54:56 | gmann | johnthetubaguy: thanks | |
| 15:08:34 | openstackgerrit | Merged openstack/os-vif master: Add ``HostPortProfileInfo`` class https://review.openstack.org/441590 | |
| 15:11:40 | bauzas | pooja: you can technically run multiple scheduler services, but since the in-memory state of the objects isn't shared between all schedulers, you can face race conditions at limits, when your cloud capacity is close to be full | |
| 15:13:19 | bauzas | pooja: that is currently being tackled by the fact the scheduler (since Ocata) now uses a Placement API service that is giving it a shared view of the state of the cloud, but that's only for a couple of resource classes (RAM and CPU, disk as well but still incorrectly reported if disks are shared between computes) | |
| 15:16:27 | melwitt | bauzas: I thought claims in the scheduler made it so running multiple schedulers won't reschedule because of different in-memory states? | |
| 15:17:02 | bauzas | melwitt: for CPU, RAM and disk, yes | |
| 15:25:49 | smcginnis | Video link from John explaining new attach https://www.youtube.com/watch?v=mrgPt0c3cUw | |
| 15:26:12 | openstackgerrit | Ed Leafe proposed openstack/nova-specs master: Return Selection Objects https://review.openstack.org/498830 | |
| 15:30:45 | mriedem | sdague: replied to your -1 in https://review.openstack.org/#/c/493323/ | |
| 15:34:41 | mriedem | sdague: the way the code series goes, everything is keyed off whether or not the bdm.attachment_id is set, and that's only ever set for a *new* attachment using the new flow, which is the very last patch in the series, | |
| 15:34:54 | mriedem | and that doesn't turn on until (1) all computes are upgraded and (2) cinder 3.44 is available | |
| 15:35:01 | mriedem | so people can roll upgrade to this functionality | |
| 15:35:16 | sdague | https://github.com/openstack/nova/blob/cfdec41eeec5fab220702efefdaafc45559aeb14/nova/compute/api.py#L3741-L3745 that's not called now? | |
| 15:35:32 | mriedem | no, because if bdm.attachment_id is None: will be True | |
| 15:36:18 | sdague | ok | |
| 15:36:35 | mriedem | sdague: https://review.openstack.org/#/c/330285/136/nova/compute/api.py@3581 | |
| 15:36:49 | sdague | that seems dangerous to be implicitly working like this, vs some real const to ensure we don't go down this path | |
| 15:38:22 | mriedem | we do the min compute service version check like this everywhere | |
| 15:39:49 | sdague | I guess I'd feel more comfortable with if bdm.version >= 3.44 instead of if bdm.attachment_id == None | |
| 15:39:52 | sdague | conceptually | |
| 15:40:18 | sdague | anyway, that's fine. It's just not as obvious as I'd ideally like on it | |
| 15:40:45 | bauzas | edmondsw_: so, about https://review.openstack.org/#/c/422696/ I checked the guidelines we have for Nova | |
| 15:41:10 | bauzas | edmondsw_: https://docs.openstack.org/nova/latest/contributor/microversions.html#f2 | |
| 15:42:00 | bauzas | edmondsw_: changing from a unclear 400 to a clearer 401 looks like legit as not requiring a microversion | |
| 15:43:24 | bauzas | sdague: am I correct? changing from 400 to 401 doesn't require a microversion, right? | |
| 15:43:31 | bauzas | alex_xu: ^ | |
| 15:46:49 | openstackgerrit | Matt Riedemann proposed openstack/nova master: DNM: test new style volume attach with live migration https://review.openstack.org/481290 | |
| 15:47:48 | sdague | bauzas: honestly, we most don't try to change those things | |
| 15:48:17 | edmondsw_ | bauzas that doesn't talk about 401. And it isn't consistent with http://specs.openstack.org/openstack/api-wg/guidelines/api_interoperability.html#evaluating-api-changes | |
| 15:49:02 | edmondsw_ | sdague we should definitely update one or the other of what bauzas linked and what I linked to be consistent... | |
| 15:49:55 | edleafe | edmondsw_: bauzas: generally, changing a 400 to a 401 would require a microversion bump. | |
| 15:49:57 | sdague | I've got to think through that neutron patch and understand | |
| 15:49:59 | bauzas | sdague: what would you recommend ? | |
| 15:50:04 | bauzas | okay | |
| 15:50:09 | bauzas | thanks | |
| 15:50:17 | sdague | I don't have the brain power in discussion rooms to think that one through right now | |
| 15:50:22 | bauzas | np | |
| 15:51:21 | bauzas | edleafe: well, I understanding the reasoning that API consumers need a programmatical way to identify what error they got and do things accordingly | |
| 15:51:41 | sdague | I honestly don't think that 401 is generically understood to be token expired | |
| 15:52:01 | sdague | which puts this into the category of churn change | |
| 15:52:21 | bauzas | the problem with the error code we returned was that it was implying some misconfig | |
| 15:52:44 | edmondsw | sdague definely doesn't mean token expired... it means unauthorized, which could be token expired but could also be token was never valid | |
| 15:53:04 | sdague | edmondsw: right, which means that the change doesn't give any more ability to automatically recover | |
| 15:53:13 | sdague | which is the justification for the change | |
| 15:53:20 | edmondsw | but in a path like this, we know it was valid originally, so we know the change is that it has expired | |
| 15:53:37 | edmondsw | sdague no, because of what I just added | |
| 15:53:54 | edleafe | edmondsw: right, the general pattern for 401 is to retry once, and if that fails, give up | |
| 15:54:11 | edleafe | (retry auth, not the original call) | |
| 15:54:20 | edmondsw | edleafe right | |
| 15:54:21 | sdague | edleafe: can you point me at 2 examples of that being the pattern in existing software | |
| 15:54:45 | edmondsw | sdague we use to have examples of that all over the place. A lot of them have now moved into keystoneauth | |
| 15:54:51 | edleafe | sdague: in SDK land it is very common | |
| 15:55:02 | sdague | edleafe: ok, 2 references with git urls please | |
| 15:55:34 | sdague | I also think the real fix here is the service role so this doesn't happen | |
| 15:55:36 | bauzas | I'm not an API specialist, but say I get a 400, I know I fscked somewhere | |
| 15:55:59 | bauzas | if I'm getting a 401, I wonder why my creds are wrong | |
| 15:56:09 | sdague | if it is claimed that this is common behavior, that's fine, I'm happy to accept them with references in code showing that it is | |
| 15:56:30 | bauzas | so maybe it's not clear that a 401 could mean something "please retry" so I do understand sdague | |
| 15:57:05 | bauzas | in general, 401 would mean full stop to me | |
| 15:57:15 | bauzas | and me checking if my creds are good | |
| 15:57:25 | sdague | the neutron interaction pattern is confusing because we sometimes use the admin in disk creds, and sometimes the user token | |
| 15:58:35 | edmondsw | sdague one e.g. that should no longer be necessary because the logic has moved into keystoneauth is https://review.openstack.org/#/c/502382/ | |
| 15:59:20 | edmondsw | bauzas yes, but if you were using token x successfully for a while, and then you make another request and get an API, you know it isn't a creds issue... a working token has stopped working, i.e. expired | |
| 15:59:47 | edmondsw | s/API/401/ | |
| 16:00:00 | bauzas | I'm looking at https://tools.ietf.org/html/rfc7235#section-3.1 and I'm puzzled :) | |
| 16:00:03 | sdague | edmondsw: you would also hit that if the user was revoked right? | |
| 16:00:38 | edmondsw | sdague if we really had revocation, yeah... and that's why you'd only retry auth once | |
| 16:01:10 | sdague | bauzas: yeh, the problem auth in HTTP means something very specific, which is not what we do | |
| 16:01:22 | sdague | so our whole thing is fudgy at best | |
| 16:01:24 | edmondsw | if reauth fails, give up... but try reauth once before you give up, because most of the time that will work and it helps so much | |
| 16:01:50 | bauzas | yeah and even microversioning a 400>401 wouldn't help the problem honestly :( | |
| 16:02:00 | sdague | edmondsw: sure, but we could also actually solve the real issue and use the service user right? | |
| 16:02:03 | edmondsw | if we never retried auth on 401, things would break randomly so often that openstack wouldn't be usable | |
| 16:02:12 | sdague | so that the expiration didn't happen | |
| 16:02:32 | edmondsw | this is so fundamental that it's baked into keystoneauth now | |
| 16:02:35 | bauzas | we could | |
| 16:02:57 | sdague | edmondsw: sure, I'm trying to go back to the actual bug and figure out how to not error at all there | |
| 16:03:06 | edleafe | sdague: Here's the one I created. If it's really important I can search other SDKs later: https://github.com/EdLeafe/pyrax/blob/master/pyrax/client.py#L224 | |
| 16:03:49 | sdague | edleafe: it would be good if you could, it's nice to see if this is a pattern within folks that write openstack projects also lives in the larger ecosystem | |
| 16:03:53 | bauzas | is there any reason why we need to pass the user token to neutron and not the service user token ? | |
| 16:04:04 | edmondsw | sdague you might be onto something with service user... I'll have to think through this case again | |
| 16:04:27 | mriedem | johnthetubaguy: L266 https://etherpad.openstack.org/p/cinder-ptg-queens | |
| 16:04:56 | sdague | because if we can not fail at all, but do the right thing, that's better than changing the error code and making all clients have to do this as a retry loop | |
| 16:05:03 | sdague | even if it's a more common retry pattern | |
| 16:05:04 | edmondsw | bauzas sdague I thought we already used the service token for all calls to neutron | |
| 16:05:12 | sdague | edmondsw: we definitely don't | |
| 16:05:15 | edmondsw | ugh | |
| 16:05:24 | mriedem | mikal: hi! | |
| 16:05:25 | sdague | because sometimes you need to operator as the user | |
| 16:05:28 | mikal | mriedem: https://blueprints.launchpad.net/nova/+spec/hurrah-for-privsep | |
| 16:05:28 | edmondsw | that does explain why service token didn't just make this work, then | |
| 16:05:47 | sdague | and I think service token wrapping may not have been added for neutron | |
| 16:05:53 | mikal | And I'm doing the context squashing now, its gonna be a couple of bigish (but very mechanical) patches | |
| 16:10:34 | edleafe | sdague: https://github.com/rackspace/gophercloud/blob/master/provider_client.go#L197-L214 | |
| 16:11:14 | edleafe | sdague: the applications created with SDKs can be long-running, and as a result token expiration is something that needs to be handled smoothly | |
| 16:13:42 | sdague | edleafe: cool, thanks | |