Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-27
18:39:40 mriedem i was like, wtf
18:39:43 dansmith yes
18:39:47 Tengu but the image corruption is the best hint for now. Have to check why - the ceph cluster isn't a cluster for now and we have some failed disks on it, so it can explain a lot. it's not in prod for now, this also explain some issues
18:41:21 mriedem dansmith: throw that in https://etherpad.openstack.org/p/nova-instance-list somewhere so we don't lose it
18:42:06 stvnoyes hi mriedem, if you get a change to re-review https://review.openstack.org/#/c/463987/ it would be great. I am on vacation next week so there's still some time this week for me to turn the review around again if it's needed. thanks.
18:43:48 mriedem ok
18:44:00 mriedem dansmith: totally unrelated, but i'm think about throwing the ceph job in the experimental queue http://tinyurl.com/ydy3jek9
18:44:18 stvnoyes johnthetubaguy: pls take a look at https://review.openstack.org/#/c/506805/ when you get a chance. it's a pretty small change, and it's needed for the cinder v3 live migrate change. thanks.
18:44:23 dansmith melwitt: ^
18:44:54 melwitt gdi
18:46:39 melwitt it would take me awhile to unroll what is going on with that job
18:46:52 mriedem this is compared to the normal dsvm tempest job http://tinyurl.com/ydemspkl
18:47:46 melwitt yeah, hm. so it was tracking okay until around the 16th
18:48:14 melwitt I'll dig into it
18:48:19 mriedem well, there were also spikes in the normal job then too, just not as bad
18:48:25 mriedem 9/23 is where it goes nuts
18:49:10 melwitt yeah. last we discussed it was at the last PTG and jbernard had some TODOs but I didn't know details about what they were. I thought the first thing was something to do with a job timeout being too short
18:49:32 mriedem he was going to start restricting the tests
18:49:47 melwitt restricting in what way?
18:49:56 mriedem L138 https://etherpad.openstack.org/p/nova-ptg-pike
18:50:55 melwitt okay. so that would be my starting point, aside from the latest craziness. which might just be more timeout and OOM stuff (have to dig)
18:51:10 efried mriedem Re-request tough love on https://review.openstack.org/#/c/488137/ please
18:51:30 mriedem melwitt: https://github.com/openstack/nova/commit/980d0fcd75c2b15ccb0af857a9848031919c6c7d merged on the 22nd
18:51:43 mriedem cinder.tests.tempest.api.volume.test_volume_revert.VolumeRevertTests.test_volume_revert_to_snapshot_after_extended is what i see failing
18:51:51 mriedem so my guess is, the ceph job doesn't care for live snapshots
18:52:19 melwitt thanks for that info
18:54:19 mriedem although the test that's failing is a cinder api test
18:54:27 mriedem and i don't see any related errors in n-cpu
18:55:52 openstackgerrit Dan Smith proposed openstack/nova master: Fix policy check performance in 2.47 https://review.openstack.org/507948
18:59:40 dansmith oops, didn't finish the commit message on that one
18:59:48 dansmith I'm just testing it in my devstack rig anyway
19:02:46 openstackgerrit Matt Riedemann proposed openstack/nova master: doc: make host aggregates examples more discoverable https://review.openstack.org/507950
19:02:50 mriedem not even the bug link
19:02:53 mriedem you were so excited
19:02:59 mriedem about getting punched in the naughty parts
19:03:07 mriedem Tengu: melwitt: https://review.openstack.org/#/c/507950/
19:03:13 mriedem ^ should help a bit with doc discovery
19:03:51 Tengu mriedem: \o/ thanks !
19:04:07 Tengu for now I'm digging in glance, as apparently it's crashed.
19:07:42 dansmith sdague: mriedem: cfriesen_: https://imgur.com/a/IQ0Vh
19:07:58 mriedem dansmith: awesome
19:08:03 mriedem also,
19:08:30 mriedem i ran the 2.53 microversion, GET /servers/detail thing again w/o your patch, to see why i had just a big difference in numbers, and you're right, it's the vm
19:08:43 mriedem so w/o your patch, it's still closer to with your patch,
19:08:51 mriedem and over half of what it was the other day on the other vm
19:08:58 mriedem so just need to chalk that up to public cloud
19:09:09 cdent what happened at 2.47?
19:09:21 mriedem cdent: we started checking policy per instance when listing instances
19:09:21 dansmith mriedem: cool
19:09:29 mriedem which adds up when you're listing 1000 instances
19:09:30 cdent ouch
19:10:34 openstackgerrit Dan Smith proposed openstack/nova master: Fix policy check performance in 2.47+ https://review.openstack.org/507948
19:11:48 cfriesen_ cdent: my bad, I didn't realize policy check was expensive
19:12:24 mriedem i'm sure i approved the change so don't worry about it
19:12:28 cdent cfriesen_: a reasonable thing to assume in a reasonable universe, but we probably left that one long ago
19:12:32 sdague I kind of wonder if there are other places with embedded policy checks like that are expensive
19:13:17 sdague cfriesen_: there is an implicit fstat because policy is live reread
19:13:53 cdent speaking of, that’s a potential next microoptimization in the unit tests. that file gets read over and over and over over and over and over and ...
19:14:44 sdague honestly, it might behoove us to change that behavior entirely, as we've got the hup handler now
19:15:16 bauzas dansmith: you trampled me
19:15:27 dansmith bauzas: I did?
19:15:35 bauzas dansmith: with Twitter
19:15:47 bauzas :p
19:16:09 bauzas so, maybe you should be the next US president given you use Twitter for trampling folks :p
19:16:23 bauzas mmm, maybe "trample" is not the right verb
19:16:29 dansmith I'm not sure what trampling I did, but I definitely need not be president
19:16:41 penick too late i'm writing you in
19:16:58 bauzas I mean, I chilled :p
19:17:03 mriedem you made sylvain spit out his coffee
19:17:08 mriedem you "floored" him
19:17:35 bauzas when I saw the tweet for 2.47 :p
19:17:45 bauzas sorry for "trampling"
19:18:19 dansmith bauzas: okay I replied to you about two seconds before you pinged me here so I thought you meant my reply was rude in some way
19:18:33 bauzas emacron: maybe you should ask French folks to stop using French but rather English ?
19:19:05 bauzas dansmith: sorry, the verb wasn't good :)
19:19:11 dansmith ack
19:19:33 bauzas "chilling" is better
19:20:36 bauzas dansmith: anyway, thanks for your tweet
19:25:58 cfriesen_ dansmith: reviewing your patch. I assume the version check is a performance optimization to avoid the policy check if we can?
19:28:02 mriedem cfriesen_: it's because we only ever care about showing flavor extra specs if you're requesting 2.47 or above
19:28:09 mriedem so don't even make the policy check otherwise
19:28:14 dansmith cfriesen_: yeah
19:28:39 sdague dansmith: so one thing to consider on that test, there is nothing in that test asserting the server list is > 1 right now
19:28:50 sdague because it's all common setup
19:28:54 dansmith sdague: true, but I did check that its 4
19:29:02 dansmith I can add another
19:29:03 sdague I thought it was 5
19:29:08 sdague I was just running it
19:29:11 dansmith it was 4
19:29:40 mriedem so just self.assertGreater(len(instances), 1) ?
19:29:54 dansmith oh no,
19:29:58 dansmith top index was 4
19:29:59 dansmith 0-4
19:30:40 sdague yeh, something like that
19:31:02 sdague just so the test is more concisely valid
19:31:43 sdague the reset seems fine to me
19:31:52 dansmith done
19:34:22 openstackgerrit Dan Smith proposed openstack/nova master: Fix policy check performance in 2.47+ https://review.openstack.org/507948
19:36:55 cfriesen_ just throwing this out there...could we use a global variable for show_extra_specs such that it's None for the first instance and then the calculated value is used for subsequent ones? That'd avoid the API changes, but globals are icky.
19:37:17 mriedem this isn't an api change

Earlier   Later