Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-06
13:37:22 bauzas oop
13:37:28 bauzas jaypipes: https://wiki.openstack.org/wiki/OpenAttestation
13:39:08 bauzas oh interesting
13:39:17 bauzas OAT has been superseded by OpenCIT
13:39:18 mriedem gmann: ok approved
13:39:20 mriedem thanks for following up
13:40:00 mriedem bauzas: jaypipes: i'm +2 on dan's migration allocations spec if one of you wants to take a gander https://review.openstack.org/#/c/498510/
13:40:07 bauzas AFAICS, instead of querying a manifest, the compute agent sends a report to the attestation server
13:40:14 bauzas mriedem: for sure
13:40:26 jaypipes alex_xu: no, I'm saying that the "trust" part of all of this stuff would be dependent on how much you trust the caller to the placement API that is setting traits. Obviously, the Placement API doesn't verify callers or attest to their content.
13:40:27 bauzas my review bag is pretty empty those days
13:41:13 bauzas jaypipes: well, the placement API accepts credentials, so the real problem is whether they accept Keystone tokens as safe enough or not
13:41:25 openstackgerrit Stephen Finucane proposed openstack/nova master: docs: Rename cellsv2_layout -> cellsv2-layout https://review.openstack.org/498821
13:41:26 openstackgerrit Stephen Finucane proposed openstack/nova master: WIP! doc: Add contents page https://review.openstack.org/498820
13:41:26 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Cleanup of existing index pages https://review.openstack.org/498819
13:41:27 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Add configuration index page https://review.openstack.org/498818
13:41:27 jaypipes bauzas: right...
13:41:27 openstackgerrit Stephen Finucane proposed openstack/nova master: doc: Add user index page https://review.openstack.org/498817
13:41:31 bauzas jaypipes: but again, it's not my problem :)
13:41:33 stephenfin ralonsoh: You about. Question about binding profiles
13:41:47 jaypipes damn you stephenfin :) was just about to +W the bottom of that.
13:42:50 mriedem bauzas: if your review bag is empty, this bug fix and the changes below it are needed for pike https://review.openstack.org/#/c/499878/
13:42:56 mriedem once that's done and backports are merged i'll cut a release
13:43:12 alex_xu jaypipes: yea, i see that, so we don't have that problem for the existed trusted filter
13:43:13 stephenfin jaypipes: You still can - it was a rebase to move the dodgy contents patch to the end
13:43:13 bauzas mriedem: I'm already on https://review.openstack.org/#/q/topic:bp/request-spec-use-by-compute
13:43:15 stephenfin :)
13:43:35 bauzas jaypipes: mriedem: dansmith: I'm gonna ask a question for https://review.openstack.org/#/c/498510/5/specs/queens/approved/migration-allocations.rst
13:43:37 mriedem bauzas: cool - gerrit somehow auto-changed my topic branch which was annoying there
13:44:37 bauzas jaypipes: mriedem: dansmith: if we consider that allocations need to be set/corrected by the conductor for migrations (and we already do that), could we consider having a single boot request to have the simple allocation claim to be also done in the conductor ?
13:44:48 bauzas I know it has been a long conversation previously
13:45:00 bauzas and I don't want to open wounds
13:45:20 bauzas but I just want to consider how that could help us having a single driver using Placement
13:45:34 alex_xu jaypipes: is it another layer problem, since you can attach traits to the resource provider, basically you already get the admin user in the system...
13:46:27 jaypipes alex_xu: I'm not saying it's a problem, per se. I'm just saying that the OpenCTI team should be aware of the fact that communication with the placement API itself is not attested.
13:46:50 openstackgerrit Stephen Finucane proposed openstack/nova master: Make eventlet hub use a monotonic clock https://review.openstack.org/434327
13:47:04 alex_xu jaypipes: yea, i see, so let me tell the opencit team...
13:47:32 bauzas jaypipes: alex_xu: tbh, the current implementation of the TrustedFilter is already considering that the communication to the OpenCIT system is not having a man in the middle
13:47:53 bauzas or that what is passed by the OpenCIT server is correct
13:48:17 bauzas that's just adding a proxy, but the design is the same
13:48:19 jaypipes bauzas: k. but this is opening up another communication channel. that's all I was saying.
13:48:46 bauzas jaypipes: sure, I'm just trying to explain that given the current implementation, those folks don't see that as a problem already :p
13:48:58 jaypipes bauzas: cool.
13:49:57 alex_xu bauzas: yea
13:50:15 mriedem bauzas: the only reason conductor is setting/correcting allocations during a move operation is because of bugs
13:50:41 bauzas I know :(
13:50:52 mriedem bauzas: well, or failures in certain cases, like live migration pre-check failing after we've already claimed on the dest host in the scheduler
13:50:57 mriedem but the force cases are bugs,
13:51:06 bauzas but since the scheduler doesn't know whether it's a move or a boot, we need to do something like that
13:51:09 mriedem and i plan on having the force scenarios still call the scheduler but with a skip_filters flag
13:51:29 mriedem the scheduler, since pike, determines if it's a move by checking for existing allocations on another node
13:51:33 mriedem and doubles those up
13:51:39 bauzas honestly, dansmith proposed to have the scheduler code to be run by the conductor service, and I just think all of this tends to that
13:51:41 mriedem the issue with force is that we bypass the scheduler altogether
13:52:13 bauzas yup
13:52:22 mriedem so are you asking if we should stop doing the claim in the scheduler and move that to conductor now in queens?
13:53:09 bauzas mriedem: yep, I'm wondering if we should reconsider the opportunity to have claims done by the conductor for the reasons you mentioned
13:53:26 bauzas #1 if you force, then you need to reconcile claims
13:53:39 mriedem i think we can still do force and have the scheduler handle the claim
13:53:45 mriedem i'm going to write a bp for that
13:53:54 mriedem since it's going to be an rpc change and require some consideration
13:53:57 bauzas #2 for move operations, there are a list of corner cases depending on the success of the move that require the conductor to reconcile allocations as well
13:54:16 mriedem sure, but for #2 we can also fail on the compute and need to cleanup allocations,
13:54:24 bauzas good point
13:54:30 mriedem which we're seeing with prep_resize failing, evacuate moveclaim failing, and unshelve claim failing
13:54:44 mriedem i do'nt really see this any different from cleaning up ports and volume attachments
13:54:52 mriedem which happen both in conductor sometimes and in the compute
13:54:59 bauzas about the force thingy, what if placement returns "Sorry"' to the scheduler if it claims ?
13:55:12 mriedem then we fail
13:55:21 bauzas I'm not sure operators would accept that
13:55:26 mriedem they are going to have to
13:55:30 bauzas LOL
13:55:50 mriedem if you force to a compute today, the claim on the compute could still fail
13:56:00 mriedem except live migration doesn't do a claim
13:56:07 mriedem but evacuate does
13:56:15 bauzas that's more complicated :
13:56:22 mriedem maybe those claims never fail because we pass an empty limits dict
13:56:33 mriedem so the claim just considers unlimited cpu/ram/disk
13:56:37 bauzas since we don't call the scheduler, the scheduler isn't passing limits to the conductor which eventually gives them to the compute
13:56:51 bauzas so we don't really verify the resource usage
13:57:00 mriedem yeah i just said the same thing
13:57:13 mriedem honestly i didn't realize that's how things worked until last week when writing some functional tests for this
13:57:15 bauzas you typed too fast, damn you
13:57:15 jaypipes mriedem, bauzas, alex_xu: https://twitter.com/jaypipes/status/905429571012616192
13:57:44 mriedem jaypipes: i can't tell if they look freaked out like that because of the storm or if that's just normal
13:57:45 bauzas jaypipes: orly? :( you're going to be impacted ? :(
13:58:06 jaypipes bauzas: yeah. we're probably going to evacuate tonight or tomorrow morning.
13:58:11 bauzas jaypipes: le woof is hugging them
13:58:13 jaypipes mriedem: just normal.
13:58:29 jaypipes heh
13:58:32 mriedem bauzas: point is, we need the resource allocation representation in placement to be accurate,
13:58:40 mriedem and since the RT isn't adjusting claims anymore,
13:58:45 mriedem after the fact i mean,
13:58:53 mriedem we can't just not create the allocations,
13:59:35 mriedem with the force stuff before with claims, the RT would still report the usage on the compute node to the scheduler and if it was full then that compute node is out of the running for placement decisions
13:59:47 mriedem if we don't report the allocations to placement, then things are incorrect for the scheduler
13:59:53 mriedem regardless of force
14:00:04 mriedem plus, as i said in the ML, i think force is a bad idea anyway
14:00:27 mriedem i'm semi ok with continuing to ignore the filters (even though the live migration force scenario still checks ram and compute filters, but in conductor code)
14:00:34 mriedem but i'm not ok with bypassing placement

Earlier   Later