| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-29 | |||
| 15:13:07 | superdan | it's just a matter of having useful data, and as you saw, it's hard to compare two runs | |
| 15:13:21 | superdan | even if not voting, if you can't actually draw conclusions... | |
| 15:13:25 | mriedem | was thinking experimental queue so it's on-demand for things we know might impact performance | |
| 15:13:51 | mriedem | my thought was the job pulls master, runs some baselines, then applies the change, runs the same tests and compares for the relative difference | |
| 15:13:51 | superdan | you can compare normalized metrics like number of db queries or something, but runtime and cpu usage are not really doable without dedicated hardware | |
| 15:14:08 | superdan | noisy neighbor problems will still skew those | |
| 15:14:17 | superdan | the window is smaller, granted, but.. | |
| 15:14:20 | mriedem | ok, so maybe a requirement is it runs on baremetal? | |
| 15:14:33 | superdan | that would be better yeah | |
| 15:14:39 | mriedem | point being, | |
| 15:14:43 | mriedem | there is an obvious need, | |
| 15:14:49 | superdan | they could be done on virt, but only if salt is applied to the result | |
| 15:14:53 | mriedem | so if we had requirements to start, we could maybe do something | |
| 15:15:03 | mriedem | i'm thinking super simple to start | |
| 15:15:06 | superdan | are you saying my box can't satisfy all of nova's perf testing needs? | |
| 15:15:28 | mriedem | sure, if you want to become the request inbox for every time we need that | |
| 15:15:52 | superdan | heh | |
| 15:16:44 | mriedem | melwitt: re the rally job, i'd be interested to know if (1) it still works and (2) if so, what does it do? like what benchmark does it compare against? | |
| 15:16:48 | melwitt | mriedem: fwiw, I was thinking the same thing recently. if we had some basic perf tests that timed instance list, delete, boot. the main things | |
| 15:16:51 | mriedem | andreykurilin: ^ maybe you can answer that | |
| 15:18:40 | andreykurilin | mriedem: hi! what is the question? :) | |
| 15:19:06 | andreykurilin | melwitt, mrieden: rally has a bunch of scenarios related to nova | |
| 15:19:39 | melwitt | mriedem: I don't know much about it other than I know some projects use it to have some monitor of perf. lemme see if I can find something real quick | |
| 15:20:01 | mriedem | andreykurilin: "re the rally job, i'd be interested to know if (1) it still works and (2) if so, what does it do? like what benchmark does it compare against?" | |
| 15:21:08 | andreykurilin | mriedem: so rally job can be used to check that things are working under some load and in concurrency mode. Also, you can specify SLAs for a single actions, like neutron job did https://github.com/openstack/neutron/blob/master/rally-jobs/neutron-neutron.yaml#L20-L21 | |
| 15:21:11 | melwitt | mriedem: found one with cinder, example patch https://review.openstack.org/#/c/501478 and the job http://logs.openstack.org/78/501478/1/check/gate-rally-dsvm-cinder-ubuntu-xenial-nv/e0eb0eb/ | |
| 15:22:28 | andreykurilin | melwitt: neutron has voting rally job for a year I think | |
| 15:22:52 | melwitt | okay, so neutron and cinder were the projects I was probably thinking of | |
| 15:22:53 | andreykurilin | cinder has non-voting job for years | |
| 15:23:05 | andreykurilin | manila has the rally job too | |
| 15:23:55 | mriedem | gibi: is this a co-worker of yours? http://forumtopics.openstack.org/cfp/details/50 | |
| 15:24:38 | mriedem | andreykurilin: that limit has to be pretty high doesn't it if the rally jobs are running on vms with probably wild variance | |
| 15:24:48 | mriedem | that was always the reason we didn't have a voting rally job in nova | |
| 15:25:15 | mriedem | and if the limit is really high, then the only things your catching are something that's really crazy and has gone off the rails.... | |
| 15:26:46 | andreykurilin | mriedem: `that was always the reason we didn't have a voting rally job in nova` it is wrong. There were several guys who dislike rally from the beggining (I do not want to name them) and blocked everything related | |
| 15:26:50 | mriedem | cdent: are you one of the people that are responsible for filtering these forum session proposals? | |
| 15:26:57 | cdent | no sir | |
| 15:27:28 | mriedem | andreykurilin: ok, personal vendettas aside, i think the general reason was the variance between runs | |
| 15:27:35 | cdent | but I was reading them anyway, so I thought I’d express an opinion, because why not? | |
| 15:28:06 | sean-k-mooney | leakypipes: sahid: sahid i do not want to top post over latest responce to vgpu implementation but can your responded to http://lists.openstack.org/pipermail/openstack-dev/2017-September/122702.html where i pointed out that any vgpu support we introduce cannot rely on mdev as amd do not use mdevs they use sriov | |
| 15:28:20 | bauzas | oh man, I now need to learn the Zuul language | |
| 15:28:47 | mriedem | mdev or bust | |
| 15:29:12 | sean-k-mooney | leakypipes: sahid we can use mdev but we have to also support sriov | |
| 15:29:24 | bauzas | mriedem: man, you don't imagine | |
| 15:29:48 | leakypipes | sean-k-mooney: I don't think anyone's saying we would *only* support mdev. | |
| 15:30:04 | andreykurilin | mriedem: about metrics and their processing. The SLA is a pluggable thing and you can use it in different ways. (90% percentiles, median...). But I agree that the results can look random due to the hardware, but anyway it can show bad trends, like it was in June - http://andreykurilin.me/trends/trends_gate-rally-dsvm-neutron-rally-ubuntu-xenial.html#/NeutronNetworks.list_agents | |
| 15:30:30 | leakypipes | sean-k-mooney: I think everyone's in agreement that we should enable multiple device management APIs, not tie us to one or another. | |
| 15:30:53 | andreykurilin | mriedem: also, the reason of making rally job voting in neutron was an ability to check the concurrency issues | |
| 15:31:03 | sean-k-mooney | leakypipes: yes | |
| 15:31:05 | sahid | sean-k-mooney: yes really, since the beginning the point was that, vgpus can be exposed with sriov or mdev | |
| 15:31:40 | cdent | sean-k-mooney: while you’re around, can you have a look at https://review.openstack.org/#/c/504540/ you were one of the people who had some ideas on how to do allocation candidate limiting, and there now lots of ideas on that spec, but some confusion on what things we are trying to optimize | |
| 15:31:55 | sean-k-mooney | cdent: sure opening it now | |
| 15:32:00 | cdent | thanks | |
| 15:32:07 | bauzas | leakypipes: sean-k-mooney: that's one of the reasons why the vGPU spec is saying we will only pass GPU resource classes | |
| 15:32:17 | bauzas | if libvirt uses mdev, then meh | |
| 15:32:27 | bauzas | if xen is using anything else, then meh | |
| 15:32:39 | sean-k-mooney | bauzas: its nto a libvirt vs xen thing | |
| 15:32:41 | bauzas | we shouldn't just leak out the technical details | |
| 15:32:42 | sahid | it's not libvirt who is using mdev, it's the hardware driver... | |
| 15:33:04 | sean-k-mooney | at the hardware level amd uses hardware based partitioning of the gpu via sriov | |
| 15:33:12 | bauzas | whatever the solution is, the interface with the compute manager is just a resource class | |
| 15:33:17 | sean-k-mooney | intel and nvidia do it via the diriver with mdevs | |
| 15:33:48 | bauzas | for the moment, libvirt supports virtual GPUs by mdevs but that's just a temporary thing AFAICU | |
| 15:33:51 | andreykurilin | mriedem: I agree that it is difficult to compare the results from different hardware(which we have in infra), but it is not a good time for openstack and I do not know the company who are ready to give reserved hardware for a performance ci | |
| 15:34:15 | bauzas | sean-k-mooney: yeah that's the driver which provides mdev types, libvirt just gets them | |
| 15:34:17 | sean-k-mooney | bauzas: the hypervisror did nto need modifcation to work with vgpus with sriov | |
| 15:34:51 | bauzas | sean-k-mooney: not sure I get your point :) | |
| 15:35:11 | sean-k-mooney | my point is that libvirt works with amd vgpus also | |
| 15:35:27 | sean-k-mooney | you just to a pci passthough of the sriov vf | |
| 15:35:47 | bauzas | ah okay, and amd driver uses SR-IOV for that? | |
| 15:35:57 | bauzas | the AMD driver isn't using mdevs ? | |
| 15:36:05 | sean-k-mooney | yes they do the virtualisation in hardware | |
| 15:36:09 | bauzas | I see | |
| 15:36:11 | bauzas | so, the thing is | |
| 15:36:11 | sean-k-mooney | they do not use mdevs | |
| 15:36:20 | bauzas | I see baby steps here | |
| 15:36:38 | bauzas | step 1/ make sure we have something workable without adding more tech deby | |
| 15:37:07 | bauzas | step 2/ work on the generic thingy from efried and "do the work" (c) | |
| 15:37:37 | bauzas | step 3/ migrate how we track PCI devices from the old world to the new world | |
| 15:38:02 | fried_rice | ++ | |
| 15:38:12 | bauzas | hah, Friday! | |
| 15:38:28 | bauzas | holy shit, I said many times I need a ZNC bot for me | |
| 15:38:40 | mriedem | andreykurilin: http://forumtopics.openstack.org/cfp/details/55 | |
| 15:38:43 | mriedem | superdan: melwitt: ^ | |
| 15:39:23 | melwitt | cool | |
| 15:40:27 | sean-k-mooney | bauwser: makes sense, my concern is coming for interop side in that i would hope at the api level we will not need to expose if it is an mdev or sriov based vgpu | |
| 15:40:51 | bauwser | sean-k-mooney: that's where I'm pretty firm in my mind | |
| 15:41:02 | bauwser | sean-k-mooney: outside the virt driver, we shouldn't leak out the details | |
| 15:41:23 | sean-k-mooney | bauwser: we may need to expose a vender in some form for guest driver reasons but other then that how we virtualise should be internal to nova | |
| 15:41:26 | bauwser | the virt driver will just report inventories | |
| 15:41:46 | bauwser | how those inventories will be populated will be different based on the driver | |
| 15:41:59 | bauwser | but that will be consistent | |
| 15:42:17 | openstackgerrit | Eric Fried proposed openstack/nova master: WIP: Use ksa adapter for cinder client https://review.openstack.org/508345 | |
| 15:42:41 | sean-k-mooney | sure and i think we can use traits on the provider to hanel the "i only have a nvida driver in my vm image" problem | |
| 15:42:44 | fried_rice | mordred ^ would you mind checking that I did my version discovery sanely here? | |
| 15:43:03 | bauwser | so that a flavor asking for an amount of 2 of the VGPU resource class and a trait of 'K800' will get 2 K800 vGPUs | |
| 15:43:21 | bauwser | that's where the API is | |
| 15:43:46 | bauwser | actually a good call for trying to fit the fried_rice's model with vGPUs | |
| 15:44:19 | sean-k-mooney | bauwser: perhaps yes or as image metadata where i say i can support x,y,z and the and flavor say i want y and we compare to make sure its compatible | |
| 15:45:10 | bauwser | sean-k-mooney: mmm | |