Earlier  
Posted Nick Remark
#openstack-nova - 2019-02-04
17:56:41 sean-k-mooney it realy has become a problem lately that we have 2 seperate code patchs. 1 for numa awre guess and anotehr for non numa guests and i just feel like have one code path that works for all guesst in a common way might be the way to go
17:57:10 sean-k-mooney efried: anyway that is slightly different then the question you asked
17:57:48 cdent sean-k-mooney: I sometimes wonder the opposite: we should have two diferent paths: one for simple vms that normal people use, anothe for the crazy path :)
17:58:13 sean-k-mooney cdent: the issue is we also have 2 sets of config options and flavor atributes
17:58:22 sean-k-mooney cdent: and people mix and match them
17:59:06 sean-k-mooney i think we can all agree that there should only be one way of modeling the resocues in placement
17:59:23 efried s/resources in placement/resources: in placement/
17:59:55 sean-k-mooney efried: placement does not do assignment
18:00:15 sean-k-mooney efried: so assignable resouces need to be tracked in the ResocueTracker also
18:00:16 efried at the RP level, not at the resource level, right.
18:00:29 efried Well, arguably they need to be tracked by the virt driver.
18:00:39 sean-k-mooney yes
18:01:01 efried I'm not sure there's a need for the resource tracker if all resources are tracked by the virt driver + placement.
18:01:25 sean-k-mooney well kind of. in the libvirt case the Resouce track is running in the compute aganet and is populated byt the virtdriver.
18:02:08 cdent efried: don't duck on that, fly that flag
18:02:18 sean-k-mooney efried: the virtdriver do not get there own db tables to track things so they have to use the resouce tracker to persist there tracked resources
18:03:37 efried sean-k-mooney: Oh, see, I disagree with that. The virt driver knows about its resources by looking at the VM. There shouldn't be a need to duplicate that information in an OpenStack database.
18:04:02 sean-k-mooney efried: the vm xml is not persisted acroudn vm reboots
18:04:19 efried oh, right, that stupid thing. We should fix *that* instead.
18:04:46 sean-k-mooney it could be maintianed in memory but im not sure that would work aross agent restarts
18:05:14 sean-k-mooney efried: well you just moving the state form the db to libvirt/xml or ram
18:05:48 efried Let me put it this way: Do any other virt drivers have this silly ephemeralness?
18:05:55 efried PowerVM doesn't. cdent, does VMWare?
18:06:12 efried If you reboot an ironic node, does it forget how much memory it had?
18:06:33 sean-k-mooney efried: all that info is stored in ironics db
18:06:47 cdent efried: vmware keeps track of itself, yes
18:06:52 sean-k-mooney libvirt db is the filesystem.
18:07:04 efried right, so why should the xml go away when you reboot it?
18:07:11 efried it's just a file in the file system, nah?
18:07:38 sean-k-mooney if the hsot is rebooted it does not
18:07:41 efried I may be tilting at windmills here, but that just seems goofy to me.
18:07:56 sean-k-mooney if the vm is rebooted it does and if we are recreated multipel vms in parallel it get racy
18:08:42 sean-k-mooney we could do it all in memory if we do there right locking to keep everything in sysnce or we can use a dbms
18:08:43 efried needing to keep track of the assignments in the nova database is just doomed. Because different platforms have different ways of identifying/tracking their resources. You'll never get the fields right for everyone. Unless they're just a big anonymous data blob.
18:08:58 cdent efried++
18:09:09 cdent virtdriver should have it's own interaction with placement
18:09:09 efried this is why we were in such pain in PowerVM-land trying to get "PCI" passthrough to work
18:09:19 efried and why I started on my "generic device management" crusade three years ago.
18:09:22 cdent nova should just "use" that info as required
18:09:23 sean-k-mooney *cough* numa toplogy blob *couch*
18:09:37 sean-k-mooney but yes you are right
18:09:50 sean-k-mooney anyway we could chagne that if we did a parallel impementation
18:10:04 sean-k-mooney but its non triavial
18:10:12 efried smy point exactly. If the numa topology is modeled in placement, it's the virt driver's responsibility to draw that model via update_provider_tree, in whatever way is appropriate for that virt.
18:10:41 sean-k-mooney efried: placement cant sotre everything we need it too
18:10:58 openstackgerrit Matt Riedemann proposed openstack/nova master: Follow up (#2) for the bw resource provider series https://review.openstack.org/634767
18:10:59 sean-k-mooney but but it should be able to handel hugepges entirly
18:11:31 sean-k-mooney we could model cpus but we would have to have 1 RP per cpu which is not a good design
18:12:28 efried Oh, I'm not saying placement should be exploded to track individual resources. Counts ought to be fine. The individual resources should be tracked by the virt driver. And trying to store those mappings in some generic database table is going to be problematic.
18:12:36 efried s/going to be//
18:13:06 sean-k-mooney efried: ya. its legacy tech debt
18:13:48 efried sean-k-mooney: How do we do pinning these days? Like this: https://docs.openstack.org/nova/pike/admin/cpu-topologies.html ?
18:14:00 sean-k-mooney so i think there are two parally efforts. 1.) modelign what can and should be modled in placemetn and 2.) refacorting the libvirt dirver to do assignment sanly
18:14:36 sean-k-mooney efried: we use the numa toplogy blob to store the free cpus
18:15:13 sean-k-mooney and then the virt driver caluates the guest topology can claims them in the host numa toplogy blob via the resouce tracker
18:16:34 sean-k-mooney efried: an important thing to remember is that hw:cpu_socket hw:cpu_threads and hw:cpu_cores are all refering to the virtual toplogy and have no baring on placemnt, the host or schdueling
18:17:17 sean-k-mooney efried: e.g. if the host has 1 socket you can sping up a vm with hw:cpu_sockets=8 and its fine
18:18:08 efried ugh. I know I have a lot to learn here, but I suspect much of it is going to be tribal.
18:18:24 efried if there are any documents that will give me some reasonably sane picture of how all this works, please do share.
18:18:48 sean-k-mooney efried: stephenfin has dont some presentation on this at the summit
18:18:55 sean-k-mooney or fosdem
18:20:35 openstackgerrit Rodolfo Alonso Hernandez proposed openstack/os-vif master: Add native implementation OVSDB API https://review.openstack.org/482226
18:20:37 cdent that etcd-compute toy I've built takes the position of "you can only have stuff that can be represented in placement". It flows nicely.
18:21:38 sean-k-mooney cdent: link?
18:22:01 sean-k-mooney https://github.com/cdent/etcd-compute?
18:22:02 cdent sean-k-mooney: https://github.com/cdent/etcd-compute I'm surprised I haven't already shown you this.
18:22:07 cdent aye
18:22:24 cdent big hack, but fun
18:23:20 sean-k-mooney the bigger the hack the more fun it becomes when it starts to work
18:23:50 sean-k-mooney cdent: are you using etcd to store any state
18:24:21 cdent etcd is used for basically two things:
18:24:35 cdent message the creation or destruction of a vm (via a watcher) to a target host
18:24:45 cdent store the IP of the created vm (while it exists)
18:24:51 openstackgerrit Matt Riedemann proposed openstack/nova master: Extend RequestGroup object for mapping https://review.openstack.org/619527
18:25:17 sean-k-mooney its architected similar to OVN in that respect
18:25:53 sean-k-mooney e.g. etcd is beign used to store teh desireds state and then distibuted workers are watching for changes and locally trying to achive teh desired state
18:26:19 sean-k-mooney so rather then an imperitive workflow its declaritve
18:26:22 cdent creation or destruction is based on the allocations: if they are there, create (with those resources), if not, destroy, in either case the same operation happens to placement: PUT /allocations/{uuid}
18:26:31 cdent yup
18:26:45 cdent it _feels_ really nice
18:27:05 sean-k-mooney it _feels_ a bit like k8s
18:27:25 sean-k-mooney perhaps with less yaml
18:27:39 cdent yes, less yaml
18:28:04 cdent my desire for using declarative messaging with vm creation predates my awareness of k8s
18:28:09 cdent but k8s provided the mechanism
18:28:16 cdent (awareness of the mechanism)
18:28:42 sean-k-mooney cdent: yes the mechanism was not invented by k8s but its one of the more promenet examples
18:29:17 cdent I was using similar mechanism in a tool called mcfeely back in the 90s
18:29:32 cdent mcfeely was your special delivery man
18:36:21 artom efried, looks like sean-k-mooney answered all your things?
18:36:44 sean-k-mooney artom: we went on a bit of a tangent
18:37:17 sean-k-mooney artom: but i think i answered his first question e.g. your stuff does not depend on placement and uses the ResocueTracker
18:37:33 sean-k-mooney *i.e. not e.g.
18:38:00 artom No, because let's be real, NUMA in placement is closer to a pipe dream than a real thing at this point (not dissing anyone, it's just a complicated thing to get working properly)
18:39:30 artom efried, and the resource tracker will always remain, because it covered the use case of tracking specific individual things (CPU 0, NUMA node 1), not just quantities of resources
18:39:38 sean-k-mooney artom: right and paraphasing efried main question. is nuam/cpu pinning in placments someting we can start working on/proposing specs for in train
18:39:53 artom Don't see why not, I mean the spec was up for Stein
18:40:07 artom It stalled, but there was no massive disagreement AFAICT
18:40:42 sean-k-mooney that was mainly because we had too many other things to do in stien
18:46:54 cdent hopefully when placement is more independent, it will be more clear that nova's inability to take advantage of placement is nova's problem, not placement's

Earlier   Later