| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-06 | |||
| 13:32:57 | alex_xu | bauzas: better answer than me :) | |
| 13:33:14 | bauzas | honestly, the OAT feature is nice | |
| 13:33:34 | bauzas | that's how it's implemented in Nova that is bad to me | |
| 13:33:48 | bauzas | s/bad/wrong | |
| 13:34:01 | jaypipes | bauzas: OAT? | |
| 13:34:14 | alex_xu | bauzas: yea, so I'm thinking we have placement API, we can have something report to the placement API direclty out of Nova | |
| 13:34:37 | bauzas | jaypipes: Open Attestation Service or something like that | |
| 13:35:19 | bauzas | jaypipes: that's basically some agent running on computes that verify whether the compute kernel is not modified | |
| 13:35:37 | bauzas | s/modified/tainted I'd say | |
| 13:35:43 | jaypipes | alex_xu: sure, you can do that. but then aren't we opening up the "trust" attack vector to anyone who can write a trait for a resource provider? ;) | |
| 13:36:08 | bauzas | jaypipes: that necessarly supposes that the Placement service can't be tainted | |
| 13:36:15 | jaypipes | bauzas: zactly ;) | |
| 13:36:27 | bauzas | that's an operator problem honestly :) | |
| 13:36:42 | bauzas | if they want to go that direction, fine by me | |
| 13:37:20 | bauzas | jaypipes: OpenAttestation | |
| 13:37:20 | alex_xu | jaypipes: so you mean the only right way is access the OAT server in each scheduling... | |
| 13:37:22 | bauzas | oop | |
| 13:37:28 | bauzas | jaypipes: https://wiki.openstack.org/wiki/OpenAttestation | |
| 13:39:08 | bauzas | oh interesting | |
| 13:39:17 | bauzas | OAT has been superseded by OpenCIT | |
| 13:39:18 | mriedem | gmann: ok approved | |
| 13:39:20 | mriedem | thanks for following up | |
| 13:40:00 | mriedem | bauzas: jaypipes: i'm +2 on dan's migration allocations spec if one of you wants to take a gander https://review.openstack.org/#/c/498510/ | |
| 13:40:07 | bauzas | AFAICS, instead of querying a manifest, the compute agent sends a report to the attestation server | |
| 13:40:14 | bauzas | mriedem: for sure | |
| 13:40:26 | jaypipes | alex_xu: no, I'm saying that the "trust" part of all of this stuff would be dependent on how much you trust the caller to the placement API that is setting traits. Obviously, the Placement API doesn't verify callers or attest to their content. | |
| 13:40:27 | bauzas | my review bag is pretty empty those days | |
| 13:41:13 | bauzas | jaypipes: well, the placement API accepts credentials, so the real problem is whether they accept Keystone tokens as safe enough or not | |
| 13:41:25 | openstackgerrit | Stephen Finucane proposed openstack/nova master: docs: Rename cellsv2_layout -> cellsv2-layout https://review.openstack.org/498821 | |
| 13:41:26 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Cleanup of existing index pages https://review.openstack.org/498819 | |
| 13:41:26 | openstackgerrit | Stephen Finucane proposed openstack/nova master: WIP! doc: Add contents page https://review.openstack.org/498820 | |
| 13:41:27 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Add user index page https://review.openstack.org/498817 | |
| 13:41:27 | jaypipes | bauzas: right... | |
| 13:41:27 | openstackgerrit | Stephen Finucane proposed openstack/nova master: doc: Add configuration index page https://review.openstack.org/498818 | |
| 13:41:31 | bauzas | jaypipes: but again, it's not my problem :) | |
| 13:41:33 | stephenfin | ralonsoh: You about. Question about binding profiles | |
| 13:41:47 | jaypipes | damn you stephenfin :) was just about to +W the bottom of that. | |
| 13:42:50 | mriedem | bauzas: if your review bag is empty, this bug fix and the changes below it are needed for pike https://review.openstack.org/#/c/499878/ | |
| 13:42:56 | mriedem | once that's done and backports are merged i'll cut a release | |
| 13:43:12 | alex_xu | jaypipes: yea, i see that, so we don't have that problem for the existed trusted filter | |
| 13:43:13 | bauzas | mriedem: I'm already on https://review.openstack.org/#/q/topic:bp/request-spec-use-by-compute | |
| 13:43:13 | stephenfin | jaypipes: You still can - it was a rebase to move the dodgy contents patch to the end | |
| 13:43:15 | stephenfin | :) | |
| 13:43:35 | bauzas | jaypipes: mriedem: dansmith: I'm gonna ask a question for https://review.openstack.org/#/c/498510/5/specs/queens/approved/migration-allocations.rst | |
| 13:43:37 | mriedem | bauzas: cool - gerrit somehow auto-changed my topic branch which was annoying there | |
| 13:44:37 | bauzas | jaypipes: mriedem: dansmith: if we consider that allocations need to be set/corrected by the conductor for migrations (and we already do that), could we consider having a single boot request to have the simple allocation claim to be also done in the conductor ? | |
| 13:44:48 | bauzas | I know it has been a long conversation previously | |
| 13:45:00 | bauzas | and I don't want to open wounds | |
| 13:45:20 | bauzas | but I just want to consider how that could help us having a single driver using Placement | |
| 13:45:34 | alex_xu | jaypipes: is it another layer problem, since you can attach traits to the resource provider, basically you already get the admin user in the system... | |
| 13:46:27 | jaypipes | alex_xu: I'm not saying it's a problem, per se. I'm just saying that the OpenCTI team should be aware of the fact that communication with the placement API itself is not attested. | |
| 13:46:50 | openstackgerrit | Stephen Finucane proposed openstack/nova master: Make eventlet hub use a monotonic clock https://review.openstack.org/434327 | |
| 13:47:04 | alex_xu | jaypipes: yea, i see, so let me tell the opencit team... | |
| 13:47:32 | bauzas | jaypipes: alex_xu: tbh, the current implementation of the TrustedFilter is already considering that the communication to the OpenCIT system is not having a man in the middle | |
| 13:47:53 | bauzas | or that what is passed by the OpenCIT server is correct | |
| 13:48:17 | bauzas | that's just adding a proxy, but the design is the same | |
| 13:48:19 | jaypipes | bauzas: k. but this is opening up another communication channel. that's all I was saying. | |
| 13:48:46 | bauzas | jaypipes: sure, I'm just trying to explain that given the current implementation, those folks don't see that as a problem already :p | |
| 13:48:58 | jaypipes | bauzas: cool. | |
| 13:49:57 | alex_xu | bauzas: yea | |
| 13:50:15 | mriedem | bauzas: the only reason conductor is setting/correcting allocations during a move operation is because of bugs | |
| 13:50:41 | bauzas | I know :( | |
| 13:50:52 | mriedem | bauzas: well, or failures in certain cases, like live migration pre-check failing after we've already claimed on the dest host in the scheduler | |
| 13:50:57 | mriedem | but the force cases are bugs, | |
| 13:51:06 | bauzas | but since the scheduler doesn't know whether it's a move or a boot, we need to do something like that | |
| 13:51:09 | mriedem | and i plan on having the force scenarios still call the scheduler but with a skip_filters flag | |
| 13:51:29 | mriedem | the scheduler, since pike, determines if it's a move by checking for existing allocations on another node | |
| 13:51:33 | mriedem | and doubles those up | |
| 13:51:39 | bauzas | honestly, dansmith proposed to have the scheduler code to be run by the conductor service, and I just think all of this tends to that | |
| 13:51:41 | mriedem | the issue with force is that we bypass the scheduler altogether | |
| 13:52:13 | bauzas | yup | |
| 13:52:22 | mriedem | so are you asking if we should stop doing the claim in the scheduler and move that to conductor now in queens? | |
| 13:53:09 | bauzas | mriedem: yep, I'm wondering if we should reconsider the opportunity to have claims done by the conductor for the reasons you mentioned | |
| 13:53:26 | bauzas | #1 if you force, then you need to reconcile claims | |
| 13:53:39 | mriedem | i think we can still do force and have the scheduler handle the claim | |
| 13:53:45 | mriedem | i'm going to write a bp for that | |
| 13:53:54 | mriedem | since it's going to be an rpc change and require some consideration | |
| 13:53:57 | bauzas | #2 for move operations, there are a list of corner cases depending on the success of the move that require the conductor to reconcile allocations as well | |
| 13:54:16 | mriedem | sure, but for #2 we can also fail on the compute and need to cleanup allocations, | |
| 13:54:24 | bauzas | good point | |
| 13:54:30 | mriedem | which we're seeing with prep_resize failing, evacuate moveclaim failing, and unshelve claim failing | |
| 13:54:44 | mriedem | i do'nt really see this any different from cleaning up ports and volume attachments | |
| 13:54:52 | mriedem | which happen both in conductor sometimes and in the compute | |
| 13:54:59 | bauzas | about the force thingy, what if placement returns "Sorry"' to the scheduler if it claims ? | |
| 13:55:12 | mriedem | then we fail | |
| 13:55:21 | bauzas | I'm not sure operators would accept that | |
| 13:55:26 | mriedem | they are going to have to | |
| 13:55:30 | bauzas | LOL | |
| 13:55:50 | mriedem | if you force to a compute today, the claim on the compute could still fail | |
| 13:56:00 | mriedem | except live migration doesn't do a claim | |
| 13:56:07 | mriedem | but evacuate does | |
| 13:56:15 | bauzas | that's more complicated : | |
| 13:56:22 | mriedem | maybe those claims never fail because we pass an empty limits dict | |
| 13:56:33 | mriedem | so the claim just considers unlimited cpu/ram/disk | |
| 13:56:37 | bauzas | since we don't call the scheduler, the scheduler isn't passing limits to the conductor which eventually gives them to the compute | |
| 13:56:51 | bauzas | so we don't really verify the resource usage | |
| 13:57:00 | mriedem | yeah i just said the same thing | |
| 13:57:13 | mriedem | honestly i didn't realize that's how things worked until last week when writing some functional tests for this | |
| 13:57:15 | jaypipes | mriedem, bauzas, alex_xu: https://twitter.com/jaypipes/status/905429571012616192 | |
| 13:57:15 | bauzas | you typed too fast, damn you | |