| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2019-02-04 | |||
| 18:06:47 | cdent | efried: vmware keeps track of itself, yes | |
| 18:06:52 | sean-k-mooney | libvirt db is the filesystem. | |
| 18:07:04 | efried | right, so why should the xml go away when you reboot it? | |
| 18:07:11 | efried | it's just a file in the file system, nah? | |
| 18:07:38 | sean-k-mooney | if the hsot is rebooted it does not | |
| 18:07:41 | efried | I may be tilting at windmills here, but that just seems goofy to me. | |
| 18:07:56 | sean-k-mooney | if the vm is rebooted it does and if we are recreated multipel vms in parallel it get racy | |
| 18:08:42 | sean-k-mooney | we could do it all in memory if we do there right locking to keep everything in sysnce or we can use a dbms | |
| 18:08:43 | efried | needing to keep track of the assignments in the nova database is just doomed. Because different platforms have different ways of identifying/tracking their resources. You'll never get the fields right for everyone. Unless they're just a big anonymous data blob. | |
| 18:08:58 | cdent | efried++ | |
| 18:09:09 | efried | this is why we were in such pain in PowerVM-land trying to get "PCI" passthrough to work | |
| 18:09:09 | cdent | virtdriver should have it's own interaction with placement | |
| 18:09:19 | efried | and why I started on my "generic device management" crusade three years ago. | |
| 18:09:22 | cdent | nova should just "use" that info as required | |
| 18:09:23 | sean-k-mooney | *cough* numa toplogy blob *couch* | |
| 18:09:37 | sean-k-mooney | but yes you are right | |
| 18:09:50 | sean-k-mooney | anyway we could chagne that if we did a parallel impementation | |
| 18:10:04 | sean-k-mooney | but its non triavial | |
| 18:10:12 | efried | smy point exactly. If the numa topology is modeled in placement, it's the virt driver's responsibility to draw that model via update_provider_tree, in whatever way is appropriate for that virt. | |
| 18:10:41 | sean-k-mooney | efried: placement cant sotre everything we need it too | |
| 18:10:58 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Follow up (#2) for the bw resource provider series https://review.openstack.org/634767 | |
| 18:10:59 | sean-k-mooney | but but it should be able to handel hugepges entirly | |
| 18:11:31 | sean-k-mooney | we could model cpus but we would have to have 1 RP per cpu which is not a good design | |
| 18:12:28 | efried | Oh, I'm not saying placement should be exploded to track individual resources. Counts ought to be fine. The individual resources should be tracked by the virt driver. And trying to store those mappings in some generic database table is going to be problematic. | |
| 18:12:36 | efried | s/going to be// | |
| 18:13:06 | sean-k-mooney | efried: ya. its legacy tech debt | |
| 18:13:48 | efried | sean-k-mooney: How do we do pinning these days? Like this: https://docs.openstack.org/nova/pike/admin/cpu-topologies.html ? | |
| 18:14:00 | sean-k-mooney | so i think there are two parally efforts. 1.) modelign what can and should be modled in placemetn and 2.) refacorting the libvirt dirver to do assignment sanly | |
| 18:14:36 | sean-k-mooney | efried: we use the numa toplogy blob to store the free cpus | |
| 18:15:13 | sean-k-mooney | and then the virt driver caluates the guest topology can claims them in the host numa toplogy blob via the resouce tracker | |
| 18:16:34 | sean-k-mooney | efried: an important thing to remember is that hw:cpu_socket hw:cpu_threads and hw:cpu_cores are all refering to the virtual toplogy and have no baring on placemnt, the host or schdueling | |
| 18:17:17 | sean-k-mooney | efried: e.g. if the host has 1 socket you can sping up a vm with hw:cpu_sockets=8 and its fine | |
| 18:18:08 | efried | ugh. I know I have a lot to learn here, but I suspect much of it is going to be tribal. | |
| 18:18:24 | efried | if there are any documents that will give me some reasonably sane picture of how all this works, please do share. | |
| 18:18:48 | sean-k-mooney | efried: stephenfin has dont some presentation on this at the summit | |
| 18:18:55 | sean-k-mooney | or fosdem | |
| 18:20:35 | openstackgerrit | Rodolfo Alonso Hernandez proposed openstack/os-vif master: Add native implementation OVSDB API https://review.openstack.org/482226 | |
| 18:20:37 | cdent | that etcd-compute toy I've built takes the position of "you can only have stuff that can be represented in placement". It flows nicely. | |
| 18:21:38 | sean-k-mooney | cdent: link? | |
| 18:22:01 | sean-k-mooney | https://github.com/cdent/etcd-compute? | |
| 18:22:02 | cdent | sean-k-mooney: https://github.com/cdent/etcd-compute I'm surprised I haven't already shown you this. | |
| 18:22:07 | cdent | aye | |
| 18:22:24 | cdent | big hack, but fun | |
| 18:23:20 | sean-k-mooney | the bigger the hack the more fun it becomes when it starts to work | |
| 18:23:50 | sean-k-mooney | cdent: are you using etcd to store any state | |
| 18:24:21 | cdent | etcd is used for basically two things: | |
| 18:24:35 | cdent | message the creation or destruction of a vm (via a watcher) to a target host | |
| 18:24:45 | cdent | store the IP of the created vm (while it exists) | |
| 18:24:51 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Extend RequestGroup object for mapping https://review.openstack.org/619527 | |
| 18:25:17 | sean-k-mooney | its architected similar to OVN in that respect | |
| 18:25:53 | sean-k-mooney | e.g. etcd is beign used to store teh desireds state and then distibuted workers are watching for changes and locally trying to achive teh desired state | |
| 18:26:19 | sean-k-mooney | so rather then an imperitive workflow its declaritve | |
| 18:26:22 | cdent | creation or destruction is based on the allocations: if they are there, create (with those resources), if not, destroy, in either case the same operation happens to placement: PUT /allocations/{uuid} | |
| 18:26:31 | cdent | yup | |
| 18:26:45 | cdent | it _feels_ really nice | |
| 18:27:05 | sean-k-mooney | it _feels_ a bit like k8s | |
| 18:27:25 | sean-k-mooney | perhaps with less yaml | |
| 18:27:39 | cdent | yes, less yaml | |
| 18:28:04 | cdent | my desire for using declarative messaging with vm creation predates my awareness of k8s | |
| 18:28:09 | cdent | but k8s provided the mechanism | |
| 18:28:16 | cdent | (awareness of the mechanism) | |
| 18:28:42 | sean-k-mooney | cdent: yes the mechanism was not invented by k8s but its one of the more promenet examples | |
| 18:29:17 | cdent | I was using similar mechanism in a tool called mcfeely back in the 90s | |
| 18:29:32 | cdent | mcfeely was your special delivery man | |
| 18:36:21 | artom | efried, looks like sean-k-mooney answered all your things? | |
| 18:36:44 | sean-k-mooney | artom: we went on a bit of a tangent | |
| 18:37:17 | sean-k-mooney | artom: but i think i answered his first question e.g. your stuff does not depend on placement and uses the ResocueTracker | |
| 18:37:33 | sean-k-mooney | *i.e. not e.g. | |
| 18:38:00 | artom | No, because let's be real, NUMA in placement is closer to a pipe dream than a real thing at this point (not dissing anyone, it's just a complicated thing to get working properly) | |
| 18:39:30 | artom | efried, and the resource tracker will always remain, because it covered the use case of tracking specific individual things (CPU 0, NUMA node 1), not just quantities of resources | |
| 18:39:38 | sean-k-mooney | artom: right and paraphasing efried main question. is nuam/cpu pinning in placments someting we can start working on/proposing specs for in train | |
| 18:39:53 | artom | Don't see why not, I mean the spec was up for Stein | |
| 18:40:07 | artom | It stalled, but there was no massive disagreement AFAICT | |
| 18:40:42 | sean-k-mooney | that was mainly because we had too many other things to do in stien | |
| 18:46:54 | cdent | hopefully when placement is more independent, it will be more clear that nova's inability to take advantage of placement is nova's problem, not placement's | |
| 18:47:05 | cdent | placement has been able to support a lot of this stuff for a _long_ time | |
| 18:47:22 | cdent | but working the changes through nova is complex | |
| 18:47:42 | sean-k-mooney | cdent: yes but it can support plociy that chage the resouces requried dependent on the host that is selected | |
| 18:47:43 | krasmussen | I made my first blueprint for a stopgap on subnet capacity aware scheduling via a filter that uses aggregates with keys containing subnets that are available to those hosts. Is this the correct place to ask for input on this? | |
| 18:47:49 | sean-k-mooney | *cant | |
| 18:48:08 | sean-k-mooney | and yes i know that by design :) | |
| 18:48:37 | sean-k-mooney | krasmussen: here or the nova meeting | |
| 18:48:46 | krasmussen | https://blueprints.launchpad.net/nova/+spec/ip-aware-scheduling-aggregate | |
| 18:49:09 | sean-k-mooney | krasmussen: the mailing list is also good however that is not ideal from a design point of view | |
| 18:49:31 | krasmussen | Figure if I'm going to take up time in a meeting I would prefer this at least have some eyes and I can improve the POC a bit more first. | |
| 18:49:45 | sean-k-mooney | krasmussen there is already a what of mapping host to reachable subnets | |
| 18:50:04 | krasmussen | Where is that accessible at? | |
| 18:51:34 | sean-k-mooney | krasmussen: are you aware of https://specs.openstack.org/openstack/nova-specs/specs/newton/implemented/neutron-routed-networks.html#neutron-routed-networks | |
| 18:52:35 | efried | artom: Save me looking for it, where's that stein spec for numa modeling in placement? | |
| 18:52:53 | artom | efried, https://review.openstack.org/#/c/552924/ | |
| 18:53:04 | efried | artom: And yes, partly because too much to do, partly because there wasn't a super big motivator for it yet, but partly because we couldn't settle on how we were going to model subtree affinity. | |
| 18:53:13 | sean-k-mooney | krasmussen: currently neutron reporst network segments to placement and uses placemetn aggreates to assocaite them with hosts | |
| 18:53:27 | efried | There were several ideas, but none was great, and we stopped discussing before we approached consensus. | |
| 18:55:52 | efried | artom: Ah, vaguely remembering that we did settle on some kind of compromise in that spec, will reread more thoroughly... | |
| 18:59:45 | krasmussen | I'm reading through all that now and following the rabbithole a bit. | |
| 19:00:20 | sean-k-mooney | krasmussen: not to be deter you by the way but the blueprint freeze for stien was in january so this feature would need to be proposed for stien | |
| 19:00:27 | sean-k-mooney | *train | |
| 19:01:56 | sean-k-mooney | krasmussen: as a comunity we are trying to move away from using nova filters for filtering based on resocue available intead prefering ot model such constriants in placement | |
| 19:02:56 | krasmussen | Thats fine :) just trying to get in the upstream flow here opposed to just local patches. | |
| 19:03:26 | sean-k-mooney | krasmussen: nova also trys not to model network backend info internally as such extending the host_state with segment/ip reachablity would be counter to the direction we have been movign over the last few cycles | |