Earlier  
Posted Nick Remark
#openstack-nova - 2019-02-25
21:46:12 sean-k-mooney cfriesen: is this the full series https://review.openstack.org/#/q/topic:bp/add-emulated-virtual-tpm+(status:open+OR+status:merged)
21:49:33 sean-k-mooney cfriesen: i think we still plan to merge the bwandwith based schduleing patch without migration support but in the case of vTPM livemigration will work right since the vTPM data is copied by qemu it will jsut be colde migrate that is missing?
21:49:45 cfriesen sean-k-mooney: paul-emile got moved internally, so I get to do it now. I'm ramping up, but I think there's a small piece missing to enable the functionality via image property and a piece missing for cold migration.
21:50:10 cfriesen sean-k-mooney: with our earlier implementation live migration just worked, I'm hoping that'll still be true.
21:52:06 cfriesen sean-k-mooney: crap, you're right. there's a bit missing to actually enable it.
21:52:21 cfriesen minor detail. :)
21:52:38 melwitt it's true we're planning to merge bw based scheduling without migration support but we're able to reject it via 403 error (I think is the plan that was landed on?) to be enabled later
21:52:52 sean-k-mooney right so i havent looked at this properly since looking at the spec.
21:53:47 sean-k-mooney cfriesen: so your in luck i have a dev env that i was usign to test artoms numa livemigration stuff that a knew enought version of qemu and libvirt
21:54:00 sean-k-mooney cfriesen: i can try and test this tommorw
21:54:29 melwitt would the vTPM stuff be able to do similar? we'd have to get some consensus about what to do, if so
21:54:42 cfriesen sean-k-mooney: okay, I'll try and get you something usable.
21:55:50 sean-k-mooney melwitt: we might be able to
21:56:00 cfriesen melwitt: currently the code automatically advertises vTPM support if qemu/libvirrt is new enough, but we could add a config option to disable that.
21:56:16 sean-k-mooney melwitt: we can certenly put a check in the conductor to reject the migration if we dont have the support
21:56:34 cfriesen yeah, we could do the conductor change too
21:57:01 sean-k-mooney cfriesen: the cold migration support should be fairly stright forward right
21:57:13 sean-k-mooney cfriesen: you just need to coppy the backing file for the vtpm?
21:57:24 melwitt yeah... config option for the support isn't something appealing
21:57:27 cfriesen sean-k-mooney: pretty much, yeah.
21:57:34 melwitt it wouldn't be able to reject before conductor? not in the API?
21:57:51 sean-k-mooney melwitt: we dont really want to do a virt dirver specific check in the api
21:58:16 mriedem surprise surprise everyone, resize/cold migrate is an RPC call from api to conductor
21:58:26 mriedem so in that case, conductor is the api
21:58:49 melwitt yeah, like isn't conductor also not a good place to do a virt driver specific check?
21:58:55 sean-k-mooney yes
21:59:01 sean-k-mooney i was about to say that
21:59:04 melwitt ok
21:59:18 sean-k-mooney the condoctor call can live migrate dest/source on the dirver
21:59:28 sean-k-mooney that is where the actual check woudl have to go
21:59:34 melwitt ah, ok
21:59:36 mriedem currently the numa checks happen in conductor
21:59:45 mriedem vgpus also landed in queens with a bunch of caveats https://docs.openstack.org/nova/latest/admin/virtual-gpu.html#caveats
21:59:47 mriedem for resize
21:59:50 sean-k-mooney mriedem: ya we could put the with the numa stuff
22:00:11 mriedem do those not also apply to vtpm, i.e. you can resize but it doesn't transfer your stuff, so you have to rebuild after that - which sucks
22:00:14 sean-k-mooney but numa worksi in hyperver and nova. actull vtpm technically is support in hyperv too
22:00:28 openstackgerrit Jim Rollenhagen proposed openstack/nova master: ironic: partition compute services by conductor group https://review.openstack.org/635006
22:01:14 mriedem with 2 weeks left i don't really want to have to think about hacking in a half-baked feature
22:01:19 mriedem i'd rather just defer to train
22:01:27 mriedem when you can do something with the actual server that has vtpm
22:01:41 sean-k-mooney cfriesen: ill need to check the spec but what did we say we woudl do for shelve ecrta
22:01:47 cfriesen okay. I'll try hard to get it all working then.
22:02:25 mriedem sean-k-mooney: i believe the spec punted on shelve
22:02:27 cfriesen pretty sure we said shelve *could* be supported by saving the TPM data as a glance image.
22:02:41 cfriesen but that we weren't planning on doing that
22:02:59 cfriesen since shelve is already broken for UEFI NVRAM
22:03:07 sean-k-mooney mriedem: right so we always defered part of it to after stien
22:03:25 sean-k-mooney cfriesen: right but i assume you would like to fix those edgecase in trian?
22:03:55 sean-k-mooney cfriesen: or rather you would like them to be fix but dont nessisarily want to have to be the one to fix them :)
22:03:55 cfriesen sean-k-mooney: I personally would love to. I don't think shelve is a big deal for our customers though.
22:03:59 mriedem windriver doesn't shelve so it doesn't matter to them
22:05:23 sean-k-mooney so to summerise i get that the preference would be to punt unless both cold/live migrate and resize work by the end of the week?
22:06:44 mriedem i only have personal preferences
22:07:00 mriedem and right now i feel the crushing weight of a deluge of last minute blueprints that need review right before FF
22:07:48 dansmith seriously, why are we even considering adding something we know we haven't fully implemented and thus can't move?
22:08:02 dansmith we already have an embarrassing number of "yeah this works, BUT YOU CAN NEVER MOVE IT" things, IMHO
22:08:11 sean-k-mooney dansmith: well live migration should work
22:08:26 sean-k-mooney its just cold/rezise but fair point
22:08:59 sean-k-mooney for livemigrat qemu copies the vtpm data iteslf but for cold we have to do it if i remember correctly
22:10:01 melwitt mriedem: I guess this convo is as good segue as any to something I was thinking about, FFE process. do we want to do the ML post, etherpad request, sponsor gathering process we have done in the past?
22:10:11 dansmith sean-k-mooney: so just the thing that *most* users will be able to trigger?
22:11:09 sean-k-mooney dansmith: yes. i am actully feeling that we should punt unless it works too but it also seams quite small
22:11:33 sean-k-mooney works -> is feature complete includign move operations
22:13:12 sean-k-mooney cfriesen: in anycase if you have someting that works im happy to test it with the other live migration feature im planning to test.
22:13:24 cfriesen sean-k-mooney: appreciated.
22:13:31 sean-k-mooney cfriesen: i get the feeling all of them will be punted to train m1 however
22:13:57 cfriesen :)
22:16:03 sean-k-mooney if we can get numa/sriov/cross-cell migration landed in m1 with vtpm and bandwith mover opertion that would be a.) alot of work and b.) alot of valuable features landed early/late depending on your view point
22:17:55 dansmith melwitt: what are the things that are realistically candidates for FFE?
22:20:43 mriedem jaypipes: good question on "wtf happens if the port is detached (or deleted) out of band while it has allocations"
22:20:51 mriedem jaypipes: i replied, but tl;dr - we don't handle it,
22:21:00 mriedem and it really kind of sucks that we (nova) are the ones managing that
22:21:54 sean-k-mooney if you delete it in neutron the send an event to nova which causes it to be detach form the vm
22:22:08 mriedem bzzt!
22:22:20 mriedem the nova code relies on getting the allocation information from the port's binding:profile
22:22:27 mriedem so if the port is deleted, we can't very well look it up
22:22:59 sean-k-mooney well it should be in the network info cache but yes that might be an issue
22:23:08 mriedem so either neutron needs to cleanup the allocation, or we have to store information about the allocation in the info cache
22:23:13 mriedem none of this is in the info cache
22:23:21 mriedem none of the requested resources / allocations stuff
22:23:40 mriedem i've asked gibi about that a few times, i.e. "you know if we just stored x in the info cache we could use it here rather than calling neutron"
22:24:00 sean-k-mooney i had tought the vif:port_profile was in the vif object in the info cache but maybe not. i havent looked at it in a while
22:24:38 melwitt dansmith: well, I'm not sure who would request one, but I thought maybe the detach root volume bp or volume-backed rebuild, or other smaller things like adding numa topo or server group to 'nova show'
22:24:47 mriedem sean-k-mooney: oh i guess it is, the binding:profile that is
22:24:55 mriedem i'm not sure it would be up to date...
22:25:10 mriedem melwitt: the volume backed rebuild isn't happening i don't tink
22:25:11 mriedem *think
22:25:14 sean-k-mooney we sore https://github.com/openstack/nova/blob/master/nova/network/model.py#L378-L400
22:25:19 dansmith melwitt: okay both of those seemed to be too far and too large to be FFE material to me,
22:25:21 sean-k-mooney which has the profile in it
22:25:26 dansmith but granted I've been a bit disconnected
22:25:30 mriedem and i'll be talking about root volume detach with Kevin_Zheng tonight but there are going to be issues with that as well
22:25:50 dansmith to me, FFE is for the final push for something that is almost done and I guess nothing pops out in my head at the moment as obviously fitting that description
22:25:59 mriedem sean-k-mooney: yeah i know yo'ure right, so we might be able to handle the port getting deleted there
22:26:06 mriedem sean-k-mooney: if the cache is up to date
22:26:18 melwitt looks like the FFE process is supposed to kick in the week after freeze anyway, so I guess I'm thinking about it too early https://docs.openstack.org/nova/latest/contributor/process.html#non-priority-feature-freeze
22:26:35 mriedem well how many weeks are there between FF and RC1?
22:26:47 mriedem if it's 2, there isn't really time for FFE unless like dansmith it's a low risk change that is already there
22:27:03 melwitt 2

Earlier   Later