Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-06
13:31:25 alex_xu bauzas: ok, probably I need to talk with that team what is the key thing needs to integrate with nova
13:31:26 bauzas alex_xu: I just want to make sure we don't create an in-tree module that would need a 3rd party system that is not tested in the gate :)
13:31:56 jaypipes alex_xu: who is using this openCIT technology?
13:32:01 bauzas and again, removing TrustedFilter from upstream doesn't mean operators or deployers could not use it
13:32:04 alex_xu bauzas: ok :)
13:32:28 alex_xu jaypipes: I don't know :)
13:32:37 bauzas well, I know that :)
13:32:44 bauzas Intel customers, I'd say
13:32:57 alex_xu bauzas: better answer than me :)
13:33:14 bauzas honestly, the OAT feature is nice
13:33:34 bauzas that's how it's implemented in Nova that is bad to me
13:33:48 bauzas s/bad/wrong
13:34:01 jaypipes bauzas: OAT?
13:34:14 alex_xu bauzas: yea, so I'm thinking we have placement API, we can have something report to the placement API direclty out of Nova
13:34:37 bauzas jaypipes: Open Attestation Service or something like that
13:35:19 bauzas jaypipes: that's basically some agent running on computes that verify whether the compute kernel is not modified
13:35:37 bauzas s/modified/tainted I'd say
13:35:43 jaypipes alex_xu: sure, you can do that. but then aren't we opening up the "trust" attack vector to anyone who can write a trait for a resource provider? ;)
13:36:08 bauzas jaypipes: that necessarly supposes that the Placement service can't be tainted
13:36:15 jaypipes bauzas: zactly ;)
13:36:27 bauzas that's an operator problem honestly :)
13:36:42 bauzas if they want to go that direction, fine by me
13:37:20 alex_xu jaypipes: so you mean the only right way is access the OAT server in each scheduling...
13:37:20 bauzas jaypipes: OpenAttestation
13:37:22 bauzas oop
13:37:28 bauzas jaypipes: https://wiki.openstack.org/wiki/OpenAttestation
13:39:08 bauzas oh interesting
13:39:17 bauzas OAT has been superseded by OpenCIT
13:39:18 mriedem gmann: ok approved
13:39:20 mriedem thanks for following up
13:40:00 mriedem bauzas: jaypipes: i'm +2 on dan's migration allocations spec if one of you wants to take a gander https://review.openstack.org/#/c/498510/
13:40:07 bauzas AFAICS, instead of querying a manifest, the compute agent sends a report to the attestation server
13:40:14 bauzas mriedem: for sure
13:40:26 jaypipes alex_xu: no, I'm saying that the "trust" part of all of this stuff would be dependent on how much you trust the caller to the placement API that is setting traits. Obviously, the Placement API doesn't verify callers or attest to their content.
13:40:27 bauzas my review bag is pretty empty those days
13:41:13 bauzas jaypipes: well, the placement API accepts credentials, so the real problem is whether they accept Keystone tokens as safe enough or not
13:41:25 openstackgerrit Stephen Finucane proposed openstack/nova master: docs: Rename cellsv2_layout -> cellsv2-layout https://review.openstack.org/498821
13:41:26 openstackgerrit Stephen Finucane proposed openstack/nova master: WIP! doc: Add contents page https://review.openstack.org/498820
13:41:26 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Cleanup of existing index pages https://review.openstack.org/498819
13:41:27 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Add configuration index page https://review.openstack.org/498818
13:41:27 jaypipes bauzas: right...
13:41:27 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Add user index page https://review.openstack.org/498817
13:41:31 bauzas jaypipes: but again, it's not my problem :)
13:41:33 stephenfin ralonsoh: You about. Question about binding profiles
13:41:47 jaypipes damn you stephenfin :) was just about to +W the bottom of that.
13:42:50 mriedem bauzas: if your review bag is empty, this bug fix and the changes below it are needed for pike https://review.openstack.org/#/c/499878/
13:42:56 mriedem once that's done and backports are merged i'll cut a release
13:43:12 alex_xu jaypipes: yea, i see that, so we don't have that problem for the existed trusted filter
13:43:13 stephenfin jaypipes: You still can - it was a rebase to move the dodgy contents patch to the end
13:43:13 bauzas mriedem: I'm already on https://review.openstack.org/#/q/topic:bp/request-spec-use-by-compute
13:43:15 stephenfin :)
13:43:35 bauzas jaypipes: mriedem: dansmith: I'm gonna ask a question for https://review.openstack.org/#/c/498510/5/specs/queens/approved/migration-allocations.rst
13:43:37 mriedem bauzas: cool - gerrit somehow auto-changed my topic branch which was annoying there
13:44:37 bauzas jaypipes: mriedem: dansmith: if we consider that allocations need to be set/corrected by the conductor for migrations (and we already do that), could we consider having a single boot request to have the simple allocation claim to be also done in the conductor ?
13:44:48 bauzas I know it has been a long conversation previously
13:45:00 bauzas and I don't want to open wounds
13:45:20 bauzas but I just want to consider how that could help us having a single driver using Placement
13:45:34 alex_xu jaypipes: is it another layer problem, since you can attach traits to the resource provider, basically you already get the admin user in the system...
13:46:27 jaypipes alex_xu: I'm not saying it's a problem, per se. I'm just saying that the OpenCTI team should be aware of the fact that communication with the placement API itself is not attested.
13:46:50 openstackgerrit Stephen Finucane proposed openstack/nova master: Make eventlet hub use a monotonic clock https://review.openstack.org/434327
13:47:04 alex_xu jaypipes: yea, i see, so let me tell the opencit team...
13:47:32 bauzas jaypipes: alex_xu: tbh, the current implementation of the TrustedFilter is already considering that the communication to the OpenCIT system is not having a man in the middle
13:47:53 bauzas or that what is passed by the OpenCIT server is correct
13:48:17 bauzas that's just adding a proxy, but the design is the same
13:48:19 jaypipes bauzas: k. but this is opening up another communication channel. that's all I was saying.
13:48:46 bauzas jaypipes: sure, I'm just trying to explain that given the current implementation, those folks don't see that as a problem already :p
13:48:58 jaypipes bauzas: cool.
13:49:57 alex_xu bauzas: yea
13:50:15 mriedem bauzas: the only reason conductor is setting/correcting allocations during a move operation is because of bugs
13:50:41 bauzas I know :(
13:50:52 mriedem bauzas: well, or failures in certain cases, like live migration pre-check failing after we've already claimed on the dest host in the scheduler
13:50:57 mriedem but the force cases are bugs,
13:51:06 bauzas but since the scheduler doesn't know whether it's a move or a boot, we need to do something like that
13:51:09 mriedem and i plan on having the force scenarios still call the scheduler but with a skip_filters flag
13:51:29 mriedem the scheduler, since pike, determines if it's a move by checking for existing allocations on another node
13:51:33 mriedem and doubles those up
13:51:39 bauzas honestly, dansmith proposed to have the scheduler code to be run by the conductor service, and I just think all of this tends to that
13:51:41 mriedem the issue with force is that we bypass the scheduler altogether
13:52:13 bauzas yup
13:52:22 mriedem so are you asking if we should stop doing the claim in the scheduler and move that to conductor now in queens?
13:53:09 bauzas mriedem: yep, I'm wondering if we should reconsider the opportunity to have claims done by the conductor for the reasons you mentioned
13:53:26 bauzas #1 if you force, then you need to reconcile claims
13:53:39 mriedem i think we can still do force and have the scheduler handle the claim
13:53:45 mriedem i'm going to write a bp for that
13:53:54 mriedem since it's going to be an rpc change and require some consideration
13:53:57 bauzas #2 for move operations, there are a list of corner cases depending on the success of the move that require the conductor to reconcile allocations as well
13:54:16 mriedem sure, but for #2 we can also fail on the compute and need to cleanup allocations,
13:54:24 bauzas good point
13:54:30 mriedem which we're seeing with prep_resize failing, evacuate moveclaim failing, and unshelve claim failing
13:54:44 mriedem i do'nt really see this any different from cleaning up ports and volume attachments
13:54:52 mriedem which happen both in conductor sometimes and in the compute
13:54:59 bauzas about the force thingy, what if placement returns "Sorry"' to the scheduler if it claims ?
13:55:12 mriedem then we fail
13:55:21 bauzas I'm not sure operators would accept that
13:55:26 mriedem they are going to have to
13:55:30 bauzas LOL
13:55:50 mriedem if you force to a compute today, the claim on the compute could still fail
13:56:00 mriedem except live migration doesn't do a claim
13:56:07 mriedem but evacuate does
13:56:15 bauzas that's more complicated :

Earlier   Later