Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-29
14:52:29 leakypipes cdent: also, nice use of the word "passel"
14:52:34 cdent leakypipes: you’re welcome. I decided this week I’d go short and focused, in part because that seemed like good timing, but also because sometimes I’m so far behind myself that I haven’t got time to do the big version. we are moving a _ton_ of code these days
14:58:15 melwitt superdan: I dunno if you saw this but I think I found a bug with the get_instance_object_sorted when there are faults. I explained it here https://review.openstack.org/#/c/505417/7/nova/tests/functional/compute/test_instance_list.py@453
14:59:50 superdan melwitt: so we're not going to merge the patch that pre-joins faults
14:59:57 superdan so that test can go away
15:00:10 superdan is there actually a problem based on how the API works now, or just that that test failed?
15:01:12 melwitt superdan: it's that _from_db_object doesn't accept pre-joined attrs, i.e. it won't set them. so even after you pre-join it lazy-loads it. and with an untargeted _context, it won't be able to and will get a None fault
15:01:16 superdan and that test was intentionally not 3 cells (see mriedem's comment above) but only because it wasn't really needed
15:01:32 superdan melwitt: okay but the api isn't asking for fault to be joined
15:02:02 superdan melwitt: the api loads the faults it wants in a batch after having collected the full list
15:02:07 melwitt superdan: you mean nova-api?
15:02:15 superdan compute/api and above but yeah
15:02:23 melwitt okay
15:02:45 superdan so you can delete that test if you want and I can re-add it if I make the fault-joining patch work better
15:03:02 superdan melwitt: this is out of the stack at the moment: https://review.openstack.org/#/c/505456/10
15:03:51 melwitt superdan: okay, so you're saying get_instance_object_sorted won't ever be used the way it is in that test
15:04:13 superdan not until I fix the above patch yeah
15:05:02 melwitt okay. yeah, I guess the only thing to look out for is if you want lazy-loads to work on the resulting cross-cell list, you'll have to make sure each of their _context are targeted to their cell
15:05:32 superdan yeah I'm not sure why they wouldn't be in this case, but it's not a scenario we request at the moment anyway
15:06:00 melwitt I think each Instance object is inited with the same untargeted context
15:06:03 superdan you could make the test use 3 cells and make it an xfail with a comment and I can come back to it when I work on the patch if you want
15:06:29 melwitt so lazy-loads wouldn't work
15:06:33 superdan melwitt: but it should be initialized with the context that scatter_gather_all_cells gave to the ...
15:06:34 superdan ohhh
15:06:35 superdan I see
15:06:44 superdan yep, I get it
15:06:47 melwitt k
15:07:12 superdan anyway, feel free to disable or remove that test for your immediate purposes and I'll circle back to it when I try to optimize the fault bit
15:07:45 openstackgerrit Chris Dent proposed openstack/nova master: [placement] gabbi tests for shared custom resource class https://review.openstack.org/485209
15:09:11 mriedem for anyone that cares, the change to get this legacy nnet job out of nova is here https://review.openstack.org/#/c/508519/
15:09:14 mriedem we're blocked until then
15:10:16 superdan bummer
15:11:48 mriedem superdan: btw, i was going to propose a forum session to brainstorm ideas for some automated perf testing like what was done this week, for a few reasons,
15:11:59 mriedem like automating it, but also find out what other projects have done, if anything,
15:12:13 superdan mriedem: we've discussed that a few times, and in boston even
15:12:21 melwitt some other projects use rally for that
15:12:28 superdan always comes back to not having stable enough stuff
15:12:38 mriedem and because there are several work group sessions, like the public cloud one, where they are talking about "what features do we want?!" and i want to say "what scale testing do we need?!"
15:12:47 mriedem this wouldn't be voting
15:13:07 superdan it's just a matter of having useful data, and as you saw, it's hard to compare two runs
15:13:21 superdan even if not voting, if you can't actually draw conclusions...
15:13:25 mriedem was thinking experimental queue so it's on-demand for things we know might impact performance
15:13:51 mriedem my thought was the job pulls master, runs some baselines, then applies the change, runs the same tests and compares for the relative difference
15:13:51 superdan you can compare normalized metrics like number of db queries or something, but runtime and cpu usage are not really doable without dedicated hardware
15:14:08 superdan noisy neighbor problems will still skew those
15:14:17 superdan the window is smaller, granted, but..
15:14:20 mriedem ok, so maybe a requirement is it runs on baremetal?
15:14:33 superdan that would be better yeah
15:14:39 mriedem point being,
15:14:43 mriedem there is an obvious need,
15:14:49 superdan they could be done on virt, but only if salt is applied to the result
15:14:53 mriedem so if we had requirements to start, we could maybe do something
15:15:03 mriedem i'm thinking super simple to start
15:15:06 superdan are you saying my box can't satisfy all of nova's perf testing needs?
15:15:28 mriedem sure, if you want to become the request inbox for every time we need that
15:15:52 superdan heh
15:16:44 mriedem melwitt: re the rally job, i'd be interested to know if (1) it still works and (2) if so, what does it do? like what benchmark does it compare against?
15:16:48 melwitt mriedem: fwiw, I was thinking the same thing recently. if we had some basic perf tests that timed instance list, delete, boot. the main things
15:16:51 mriedem andreykurilin: ^ maybe you can answer that
15:18:40 andreykurilin mriedem: hi! what is the question? :)
15:19:06 andreykurilin melwitt, mrieden: rally has a bunch of scenarios related to nova
15:19:39 melwitt mriedem: I don't know much about it other than I know some projects use it to have some monitor of perf. lemme see if I can find something real quick
15:20:01 mriedem andreykurilin: "re the rally job, i'd be interested to know if (1) it still works and (2) if so, what does it do? like what benchmark does it compare against?"
15:21:08 andreykurilin mriedem: so rally job can be used to check that things are working under some load and in concurrency mode. Also, you can specify SLAs for a single actions, like neutron job did https://github.com/openstack/neutron/blob/master/rally-jobs/neutron-neutron.yaml#L20-L21
15:21:11 melwitt mriedem: found one with cinder, example patch https://review.openstack.org/#/c/501478 and the job http://logs.openstack.org/78/501478/1/check/gate-rally-dsvm-cinder-ubuntu-xenial-nv/e0eb0eb/
15:22:28 andreykurilin melwitt: neutron has voting rally job for a year I think
15:22:52 melwitt okay, so neutron and cinder were the projects I was probably thinking of
15:22:53 andreykurilin cinder has non-voting job for years
15:23:05 andreykurilin manila has the rally job too
15:23:55 mriedem gibi: is this a co-worker of yours? http://forumtopics.openstack.org/cfp/details/50
15:24:38 mriedem andreykurilin: that limit has to be pretty high doesn't it if the rally jobs are running on vms with probably wild variance
15:24:48 mriedem that was always the reason we didn't have a voting rally job in nova
15:25:15 mriedem and if the limit is really high, then the only things your catching are something that's really crazy and has gone off the rails....
15:26:46 andreykurilin mriedem: `that was always the reason we didn't have a voting rally job in nova` it is wrong. There were several guys who dislike rally from the beggining (I do not want to name them) and blocked everything related
15:26:50 mriedem cdent: are you one of the people that are responsible for filtering these forum session proposals?
15:26:57 cdent no sir
15:27:28 mriedem andreykurilin: ok, personal vendettas aside, i think the general reason was the variance between runs
15:27:35 cdent but I was reading them anyway, so I thought I’d express an opinion, because why not?
15:28:06 sean-k-mooney leakypipes: sahid: sahid i do not want to top post over latest responce to vgpu implementation but can your responded to http://lists.openstack.org/pipermail/openstack-dev/2017-September/122702.html where i pointed out that any vgpu support we introduce cannot rely on mdev as amd do not use mdevs they use sriov
15:28:20 bauzas oh man, I now need to learn the Zuul language
15:28:47 mriedem mdev or bust
15:29:12 sean-k-mooney leakypipes: sahid we can use mdev but we have to also support sriov
15:29:24 bauzas mriedem: man, you don't imagine
15:29:48 leakypipes sean-k-mooney: I don't think anyone's saying we would *only* support mdev.
15:30:04 andreykurilin mriedem: about metrics and their processing. The SLA is a pluggable thing and you can use it in different ways. (90% percentiles, median...). But I agree that the results can look random due to the hardware, but anyway it can show bad trends, like it was in June - http://andreykurilin.me/trends/trends_gate-rally-dsvm-neutron-rally-ubuntu-xenial.html#/NeutronNetworks.list_agents
15:30:30 leakypipes sean-k-mooney: I think everyone's in agreement that we should enable multiple device management APIs, not tie us to one or another.
15:30:53 andreykurilin mriedem: also, the reason of making rally job voting in neutron was an ability to check the concurrency issues
15:31:03 sean-k-mooney leakypipes: yes
15:31:05 sahid sean-k-mooney: yes really, since the beginning the point was that, vgpus can be exposed with sriov or mdev
15:31:40 cdent sean-k-mooney: while you’re around, can you have a look at https://review.openstack.org/#/c/504540/ you were one of the people who had some ideas on how to do allocation candidate limiting, and there now lots of ideas on that spec, but some confusion on what things we are trying to optimize
15:31:55 sean-k-mooney cdent: sure opening it now
15:32:00 cdent thanks
15:32:07 bauzas leakypipes: sean-k-mooney: that's one of the reasons why the vGPU spec is saying we will only pass GPU resource classes
15:32:17 bauzas if libvirt uses mdev, then meh
15:32:27 bauzas if xen is using anything else, then meh
15:32:39 sean-k-mooney bauzas: its nto a libvirt vs xen thing
15:32:41 bauzas we shouldn't just leak out the technical details
15:32:42 sahid it's not libvirt who is using mdev, it's the hardware driver...
15:33:04 sean-k-mooney at the hardware level amd uses hardware based partitioning of the gpu via sriov
15:33:12 bauzas whatever the solution is, the interface with the compute manager is just a resource class

Earlier   Later