Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-27
18:28:28 mriedem order only matters for performance
18:28:43 Tengu and I rebooted the controllers in order to ensure all is running at the latest config version
18:28:47 melwitt Tengu: no but you will want to check nova-scheduler logs to make sure some other filter isn't rejecting it
18:28:48 Tengu mriedem: hmm ok.
18:29:12 melwitt at DEBUG log level. it's possible something else is going wrong and not the key match for the metadata
18:29:16 Tengu melwitt: yup, but I didn't see anything. the instance "directory" was created on the right node in /var/lib/nova/instances
18:29:31 melwitt if you're getting NoValidHost you should see something
18:29:49 Tengu but after a while, paff, directory is removed, and crash, "no host found"… although it actually HAD found a host
18:30:00 melwitt unless a compute host rejected the request in which case you should see an error in the nova-compute logs or the nova-conductor logs
18:30:08 Tengu hmmm.
18:30:29 Tengu will check that one.
18:30:53 melwitt the way it works is if scheduling filters all pass, it goes to nova-compute, if something fails while it builds it, it will tear it down, log stuff, and try to reschedule to another host if you have retries configured
18:31:05 Tengu what would be the patter of a rejection in nova-compute.log ?
18:31:16 Tengu hmm ok.
18:31:21 Tengu I have the retryfilter
18:31:23 melwitt should see something logged at ERROR level I think
18:31:28 melwitt in nova-compute
18:31:29 Tengu think this one will try to re-schedule
18:31:45 Tengu duh
18:32:05 Tengu corrupted image download o_O
18:32:47 Tengu that might explain a bit. but that would point the glance storage
18:33:17 openstackgerrit Eric Berglund proposed openstack/nova master: PowerVM Driver: config drive https://review.openstack.org/409404
18:33:18 openstackgerrit Eric Berglund proposed openstack/nova master: WIP(5): PowerVM driver: ovs vif https://review.openstack.org/422512
18:33:31 Tengu although… hmm. timestamp doesn't really match. will dig a bit more.
18:34:31 melwitt Tengu: yeah, so you have other issues there. but as long as the instance is always landing on the host where the aggregate meta matches the flavor, you know at least the extra specs filtering is working correctly
18:34:44 melwitt (for your original concern)
18:35:11 Tengu melwitt: right.
18:35:37 Tengu so my debug steps weren't that wrong. I should have had a better look to the nova-compute.log file though.
18:36:44 mriedem you can also trace the request id and/or instance id through the logs if you have your logs pumped to an ELK stack
18:37:05 mriedem or journald like in devstack
18:38:29 Tengu for now we don't have an ELK (it will run on the openstack… well, yes, that might cause some issues at some point ;)).
18:38:34 Tengu but we want to do that, yep.
18:39:01 dansmith mriedem: https://imgur.com/a/FY7Oq
18:39:14 dansmith mriedem: over about 300 runs, my patch is consistently faster than master
18:39:37 mriedem oh that's w/o the policy fix :)
18:39:40 mriedem i was like, wtf
18:39:43 dansmith yes
18:39:47 Tengu but the image corruption is the best hint for now. Have to check why - the ceph cluster isn't a cluster for now and we have some failed disks on it, so it can explain a lot. it's not in prod for now, this also explain some issues
18:41:21 mriedem dansmith: throw that in https://etherpad.openstack.org/p/nova-instance-list somewhere so we don't lose it
18:42:06 stvnoyes hi mriedem, if you get a change to re-review https://review.openstack.org/#/c/463987/ it would be great. I am on vacation next week so there's still some time this week for me to turn the review around again if it's needed. thanks.
18:43:48 mriedem ok
18:44:00 mriedem dansmith: totally unrelated, but i'm think about throwing the ceph job in the experimental queue http://tinyurl.com/ydy3jek9
18:44:18 stvnoyes johnthetubaguy: pls take a look at https://review.openstack.org/#/c/506805/ when you get a chance. it's a pretty small change, and it's needed for the cinder v3 live migrate change. thanks.
18:44:23 dansmith melwitt: ^
18:44:54 melwitt gdi
18:46:39 melwitt it would take me awhile to unroll what is going on with that job
18:46:52 mriedem this is compared to the normal dsvm tempest job http://tinyurl.com/ydemspkl
18:47:46 melwitt yeah, hm. so it was tracking okay until around the 16th
18:48:14 melwitt I'll dig into it
18:48:19 mriedem well, there were also spikes in the normal job then too, just not as bad
18:48:25 mriedem 9/23 is where it goes nuts
18:49:10 melwitt yeah. last we discussed it was at the last PTG and jbernard had some TODOs but I didn't know details about what they were. I thought the first thing was something to do with a job timeout being too short
18:49:32 mriedem he was going to start restricting the tests
18:49:47 melwitt restricting in what way?
18:49:56 mriedem L138 https://etherpad.openstack.org/p/nova-ptg-pike
18:50:55 melwitt okay. so that would be my starting point, aside from the latest craziness. which might just be more timeout and OOM stuff (have to dig)
18:51:10 efried mriedem Re-request tough love on https://review.openstack.org/#/c/488137/ please
18:51:30 mriedem melwitt: https://github.com/openstack/nova/commit/980d0fcd75c2b15ccb0af857a9848031919c6c7d merged on the 22nd
18:51:43 mriedem cinder.tests.tempest.api.volume.test_volume_revert.VolumeRevertTests.test_volume_revert_to_snapshot_after_extended is what i see failing
18:51:51 mriedem so my guess is, the ceph job doesn't care for live snapshots
18:52:19 melwitt thanks for that info
18:54:19 mriedem although the test that's failing is a cinder api test
18:54:27 mriedem and i don't see any related errors in n-cpu
18:55:52 openstackgerrit Dan Smith proposed openstack/nova master: Fix policy check performance in 2.47 https://review.openstack.org/507948
18:59:40 dansmith oops, didn't finish the commit message on that one
18:59:48 dansmith I'm just testing it in my devstack rig anyway
19:02:46 openstackgerrit Matt Riedemann proposed openstack/nova master: doc: make host aggregates examples more discoverable https://review.openstack.org/507950
19:02:50 mriedem not even the bug link
19:02:53 mriedem you were so excited
19:02:59 mriedem about getting punched in the naughty parts
19:03:07 mriedem Tengu: melwitt: https://review.openstack.org/#/c/507950/
19:03:13 mriedem ^ should help a bit with doc discovery
19:03:51 Tengu mriedem: \o/ thanks !
19:04:07 Tengu for now I'm digging in glance, as apparently it's crashed.
19:07:42 dansmith sdague: mriedem: cfriesen_: https://imgur.com/a/IQ0Vh
19:07:58 mriedem dansmith: awesome
19:08:03 mriedem also,
19:08:30 mriedem i ran the 2.53 microversion, GET /servers/detail thing again w/o your patch, to see why i had just a big difference in numbers, and you're right, it's the vm
19:08:43 mriedem so w/o your patch, it's still closer to with your patch,
19:08:51 mriedem and over half of what it was the other day on the other vm
19:08:58 mriedem so just need to chalk that up to public cloud
19:09:09 cdent what happened at 2.47?
19:09:21 mriedem cdent: we started checking policy per instance when listing instances
19:09:21 dansmith mriedem: cool
19:09:29 mriedem which adds up when you're listing 1000 instances
19:09:30 cdent ouch
19:10:34 openstackgerrit Dan Smith proposed openstack/nova master: Fix policy check performance in 2.47+ https://review.openstack.org/507948
19:11:48 cfriesen_ cdent: my bad, I didn't realize policy check was expensive
19:12:24 mriedem i'm sure i approved the change so don't worry about it
19:12:28 cdent cfriesen_: a reasonable thing to assume in a reasonable universe, but we probably left that one long ago
19:12:32 sdague I kind of wonder if there are other places with embedded policy checks like that are expensive
19:13:17 sdague cfriesen_: there is an implicit fstat because policy is live reread
19:13:53 cdent speaking of, that’s a potential next microoptimization in the unit tests. that file gets read over and over and over over and over and over and ...
19:14:44 sdague honestly, it might behoove us to change that behavior entirely, as we've got the hup handler now
19:15:16 bauzas dansmith: you trampled me
19:15:27 dansmith bauzas: I did?
19:15:35 bauzas dansmith: with Twitter
19:15:47 bauzas :p
19:16:09 bauzas so, maybe you should be the next US president given you use Twitter for trampling folks :p
19:16:23 bauzas mmm, maybe "trample" is not the right verb

Earlier   Later