| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-01-29 | |||
| 17:21:22 | stephenfin | jaypipes: Any chance you could take a look at this today? https://review.openstack.org/#/c/537363/ | |
| 17:24:43 | cfriesen | mriedem: is there a reason not to use a service token here? (other than that they're experimental) | |
| 17:25:43 | melwitt | mriedem: I wondered about that too (the eventlet os.open thing). I didn't realize the timing coincided with that update | |
| 17:26:05 | melwitt | *the timing of the constant fails | |
| 17:31:20 | Spazmotic | aigoo.. that's enough computers for today | |
| 17:31:24 | Spazmotic | Have a good night everyone. | |
| 17:40:33 | mriedem | cfriesen: are you conflating the service user thing here? | |
| 17:40:39 | mriedem | cfriesen: different issues | |
| 17:40:56 | mriedem | the service user thing is for re-auth if a user token expires | |
| 17:41:16 | mriedem | this is different, it's just creds to be able to have nova do stuff when there is no user token | |
| 17:41:37 | mriedem | the service user token stuff for re-auth should probably no longer be called experimental | |
| 17:41:58 | mriedem | i don't know of anyone that's done performance testing with it at scale, but haven't heard anyone complain about it either, and it's optional | |
| 17:42:34 | mriedem | cfriesen: speaking of performance, did you ever get any updates on the perf regression you reported last week/ | |
| 17:42:35 | mriedem | ? | |
| 17:43:23 | jaypipes | stephenfin: on it. | |
| 17:47:33 | melwitt | mriedem: this looks good, but are we supposed to not approve things yet because of zuul? https://review.openstack.org/#/c/538961 | |
| 17:48:17 | cfriesen | mriedem: we've got some additional data, but no smoking guns. the cpu usage just seems higher overall, with the cpu usage of most services looking more spread-out and less spiky. | |
| 17:48:25 | Spazmotic | zuul is all better ap parently | |
| 17:48:40 | Spazmotic | Was restarted and caught up | |
| 17:48:59 | Spazmotic | Was it's zuul so all better is relative.. but yeah anyway | |
| 17:49:03 | melwitt | Spazmotic: ah, thanks. I see the status alert now in the backscroll | |
| 17:51:12 | Spazmotic | np melwitt. johnthetubaguy if you can get a cahhance to review my notes on that commit today i'd appreciate it, but don't stress over it if not. If you decide to +2 it i'll see if some of the other UK guys can take a look tomorrow to finish it off :) | |
| 17:51:15 | Spazmotic | Night folks, have a good rest of your day | |
| 17:58:30 | cfriesen | mriedem: when authenticating a token, is it possible for services to explicitly check whether a valid service token was attached? (This is related to the "glance won't allow update of ceph image location during nova snapshot" issue.) | |
| 18:01:45 | mriedem | cfriesen: that sounds like a question for edmondsw or lbragstad | |
| 18:02:16 | cfriesen | mriedem: cool, will ping them | |
| 18:02:23 | edmondsw | cfriesen not sure I understood the question | |
| 18:03:04 | cfriesen | edmondsw: there's an issue (https://bugs.launchpad.net/openstack-ansible/+bug/1639940) where with ceph-backed instances/images nova tries to make a new image and then update the location later | |
| 18:03:05 | openstack | Launchpad bug 1639940 in openstack-ansible "Snapshots of instances launched from images fails with Ceph as storage." [Medium,Incomplete] - Assigned to Logan V (loganv) | |
| 18:03:25 | cfriesen | edmondsw: this used to work with glance v1, but with v2 it fails because glance doesn't allow updating the image location | |
| 18:03:52 | cfriesen | edmondsw: we were wondering whether we could use service tokens to allow glance to determine that the request is coming from another openstack service rather than a "normal" user | |
| 18:04:12 | edmondsw | ah, I see | |
| 18:04:37 | edmondsw | today, service tokens are only used re: expirations | |
| 18:04:59 | edmondsw | i.e., if the user token has expired, it'll still allow the operation if the service token isn't expired | |
| 18:05:56 | edmondsw | we have talked about going beyond that, and actually checking policy/RBAC based on service token instead of user token, but that has not been implemented | |
| 18:07:34 | cfriesen | do the APIs exist to allow glance to extract/validate the user token separately? or is that something that would need to be added in keystone? | |
| 18:07:46 | cfriesen | bah, service token, not user token | |
| 18:08:14 | edmondsw | cfriesen looking | |
| 18:08:18 | melwitt | cfriesen: it still fails after setting 'show_multiple_locations = True' as mentioned in comment 7? | |
| 18:08:40 | cfriesen | melwitt: presumably it works, but there are big security warnings about not setting that to True | |
| 18:09:10 | mriedem | cfriesen: this sounds very similar to a bug we have with glance v2 and shelve where the shelved snapshot image has a different set of project_id/user_id than the admin token that tries to get the image when unshelving, | |
| 18:09:17 | mriedem | and i thought we could use the member stuff with glancev2 | |
| 18:09:33 | cfriesen | melwitt: the current recommendation from the glance people was to have a whole separate glance node with a different config file just for nova to talk to it. | |
| 18:09:41 | mriedem | https://bugs.launchpad.net/nova/+bug/1675791 | |
| 18:09:43 | openstack | Launchpad bug 1675791 in OpenStack Compute (nova) "Instance created by demo user(non-admin), shelved by admin and unshelved by demo user --> ends up in error state " [Medium,Triaged] - Assigned to Damini Chopra (damini) | |
| 18:09:44 | melwitt | cfriesen: yeah. from what I understand, the fast clone has the caveat of the security issues | |
| 18:11:12 | melwitt | that is, you have to be in an environment where exposure of the image location urls isn't a problematic in order to use COW clone | |
| 18:11:23 | mriedem | err in the case of shelve, the elevated admin context creates the snapshot, and when the user goes to unshelve the instance, it fails b/c the non-admin user doesn't have access to the image created for them | |
| 18:12:55 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Check for leaked server resource allocations in post_test_hook https://review.openstack.org/538510 | |
| 18:12:56 | mriedem | gibi: cdent: ^ hark back to an old conversation about testing for allocation cleanups | |
| 18:15:33 | mriedem | dansmith: ^ we talked about that in denver i think | |
| 18:18:07 | dansmith | I believe you | |
| 18:18:54 | edmondsw | cfriesen looks like the context object should have info about the service_token if one was used | |
| 18:19:41 | edmondsw | though I think you're probably treading on thin ice trying to use those in a way that they were not really intended to be used | |
| 18:19:54 | edmondsw | better run what you are thinking by lbragstad | |
| 18:27:47 | efried | edleafe Be careful what you ask for. See -dev ML. | |
| 18:31:54 | lbragstad | cfriesen: edmondsw summed it up pretty well, the service token work is kind of a long road.. it sounds like you want to use service tokens to validation/determine more than just a "yes, this token is valid" or "no, it isn't"? | |
| 18:35:34 | edleafe | efried: thanks for that. So if I'm understanding things correctly, the "res.pool" is the root RP that placement would return | |
| 18:35:59 | efried | edleafe: Depends which model we're going with. | |
| 18:36:47 | efried | edleafe: That would be model (C) | |
| 18:39:15 | edleafe | efried: (C) makes the most sense to me, based on my limited familiarity | |
| 18:39:41 | efried | edleafe: But it suffers from at least two drawbacks. | |
| 18:39:43 | edleafe | efried: we don't want to fall into the trap of making everything fit the Nova model | |
| 18:40:32 | efried | edleafe: What do you mean? This *is* nova. | |
| 18:40:47 | efried | edleafe: You mean the libvirt model? | |
| 18:41:35 | edleafe | efried: yeah, like we tried to do with ironic | |
| 18:42:40 | edleafe | efried: from a placement POV, all we care about are the things that the consumer wants us to track. For generic libvirt nova, that would be compute nodes. For vmware nova, it would be resource pools | |
| 18:44:11 | efried | Dig. So we should look toward closing those gaps. In general, we should be working to make nova + placement + <virt> work smoothly, for all values of <virt>. | |
| 18:44:30 | edleafe | yeah | |
| 18:45:00 | edleafe | Placement should always be returning the thing that the consumer needs to proceed with a request | |
| 18:46:15 | efried | I need to close the time box on this. I'm going to make sure the PTG etherpad has an entry for this, and then get back to my regularly-scheduled programming. | |
| 18:46:33 | edleafe | efried: cool. Thanks for writing up that email | |
| 18:46:39 | efried | yahyoubetcha | |
| 19:00:29 | mriedem | jaypipes: melwitt: i'm going to undo the exception code handline refactor in https://review.openstack.org/#/c/538961/ - didn't mean to include that in this patch and it muddies the backpot | |
| 19:00:31 | mriedem | *backport | |
| 19:01:03 | melwitt | ack | |
| 19:04:17 | openstackgerrit | Sen Yang proposed openstack/python-novaclient master: Implement hypervisor hostname exact pattern match https://review.openstack.org/520187 | |
| 19:07:18 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Rollback instance.image_ref on failed rebuild https://review.openstack.org/538961 | |
| 19:07:19 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Collapse duplicate error handling in rebuild_instance https://review.openstack.org/539001 | |
| 19:11:53 | jaypipes | mriedem: +Wallaby'd | |
| 19:12:04 | mriedem | thanks | |
| 19:14:50 | efried | jaypipes: Can we continue the discussion as to whether or not we should in fact be requiring update_provider_tree to return True/False? | |
| 19:18:26 | openstackgerrit | Matt Riedemann proposed openstack/nova stable/pike: Rollback instance.image_ref on failed rebuild https://review.openstack.org/539003 | |
| 19:18:29 | cfriesen | lbragstad: just back from lunch. yes, there's a scenario where we would like to answer the question "did this request come from another openstack service, not a 'normal' user". was hoping to use service tokens for this. | |
| 19:20:40 | jaypipes | efried: sure | |
| 19:21:58 | efried | jaypipes: So first off, it's trivial and inexpensive for report client to figure it out, so there's no real *need* for virt to tell us. | |
| 19:22:35 | efried | jaypipes: And after talking through a couple of potential impls from VMWare, it became clear that there's certainly the possibility that it would be awkward for the virt driver to figure it out. | |
| 19:23:11 | efried | jaypipes: For example, one viable implementation is to say, "I don't care what you gave me, I'm going to delete everything and build the ProviderTree as I know it from scratch" | |
| 19:23:34 | lbragstad | cfriesen: that kinda sounds like federation? | |
| 19:23:36 | openstackgerrit | melanie witt proposed openstack/nova stable/pike: Stop globally caching host states in scheduler HostManager https://review.openstack.org/539005 | |
| 19:24:39 | efried | jaypipes: Softer than that, it's IMO a source of extra unnecessary bugs to ask the virt driver to make sure they get that bool return correct. | |
| 19:25:10 | cfriesen | lbragstad: not sure federation applies since it's all within the same cloud. We just want glance to be able to special-case a request to change an image location if it comes from nova, but not if it comes from a regular user. | |
| 19:25:31 | ameeda | gibi: now its okay, can you please merge the bug ? or it need something else ? | |
| 19:25:52 | efried | jaypipes: Put another way: why have two chunks of code doing the same thing when one will do? | |
| 19:26:31 | jaypipes | efried: ok | |
| 19:27:04 | jaypipes | efried: I just thought it would make the RT's life easier if it could say "ok, no changes from virt driver... just move on" | |
| 19:27:12 | lbragstad | cfriesen: oh - sorry, for some reason i was thinking of different services | |
| 19:27:31 | efried | jaypipes: Definitely could have worked out that way. | |
| 19:29:26 | lbragstad | cfriesen: that sounds like new territory for service tokens, the first thing we started working on with them was the whole long running operation issue.. it'd be good to sync with jamielennox though | |
| 19:30:10 | lbragstad | cfriesen: he was one of the original people driving the effort, so i wouldn't be surprised if he's ventured down a couple different paths similar to what you're describing | |