| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-27 | |||
| 18:19:33 | melwitt | I'm reading it again now | |
| 18:19:35 | mriedem | i am definitely not an expert here on the existing capabilities and gaps | |
| 18:19:36 | Tengu | melwitt: same for me - actually, for now, we're unable to start any instance because it doesn't find any host to run it | |
| 18:20:45 | Tengu | mriedem: but maybe it's "just" the metadata format that fails me. is there any doc for that? | |
| 18:20:57 | melwitt | Tengu: and you added gen1=true and gen2=true to your host aggregates? | |
| 18:21:11 | Tengu | yup, as a metadata as well | |
| 18:21:40 | mriedem | the now deleted ops guide might have had something specific for this | |
| 18:21:48 | Tengu | :'( | |
| 18:22:14 | cdent | it got moved to the wiki? | |
| 18:22:14 | Tengu | I found a doc saying the metadata on the flavor should be in the form aggregate_instance_extra_specs:gen1='true' | |
| 18:22:20 | Tengu | but that doesn't work either | |
| 18:22:20 | cfriesen_ | mriedem: dansmith: just saw the mention of microversion 2.47...there was already a call to "instance.get_flavor()" previously, so I had assumed it would get the whole flavor. I suspect you're right that it's lazy-loading extra-specs. | |
| 18:22:33 | mriedem | it would be in here if it existed https://docs.openstack.org/nova/latest/admin/index.html | |
| 18:22:40 | dansmith | cfriesen_: no it's policy | |
| 18:22:55 | dansmith | cfriesen_: the policy is checked per instance now, which is an fs call at least | |
| 18:23:09 | cfriesen_ | dansmith: ah...I had a networking glitch, missed some irc. | |
| 18:23:43 | cfriesen_ | dansmith: fix is what, cache the policy? | |
| 18:23:56 | dansmith | cfriesen_: check once per list and not once per instance | |
| 18:24:01 | dansmith | cfriesen_: I'm cooking it up now | |
| 18:24:03 | dansmith | smells like bacon | |
| 18:24:16 | cfriesen_ | dansmith: do we even need that? couldn't we check it the first time and cache it? | |
| 18:24:26 | dansmith | cfriesen_: that's what I just said | |
| 18:24:37 | cfriesen_ | I meant the first time on process startup | |
| 18:24:54 | mriedem | Tengu: i've seen a better doc than that red hat one, sec | |
| 18:25:12 | dansmith | cfriesen_: it depends per request | |
| 18:25:46 | Tengu | mriedem: that would be nice :) | |
| 18:25:47 | cfriesen_ | dansmith: ah, of course | |
| 18:25:53 | melwitt | Tengu: I found this doc https://docs.openstack.org/ocata/config-reference/compute/schedulers.html#host-aggregates | |
| 18:26:04 | dansmith | mriedem: confirmed the knee in the same place on master | |
| 18:26:05 | Tengu | I've also followed https://blog.russellbryant.net/2013/05/21/availability-zones-and-host-aggregates-in-openstack-compute-nova/ - but failed. | |
| 18:26:27 | Tengu | melwitt: ah, ocata, might work, pike is just one version ahead. will check that, thanks! | |
| 18:26:43 | mriedem | melwitt: yeah https://docs.openstack.org/ocata/config-reference/compute/schedulers.html#example-specify-compute-hosts-with-ssds | |
| 18:26:57 | mriedem | openstack flavor set --property aggregate_instance_extra_specs:ssd=true ssd.large | |
| 18:27:08 | Tengu | duh… ok, I was also on that one -.-' | |
| 18:27:17 | melwitt | Tengu: the main thing I saw ppl run into a snag is that you apparently have to use that prefix when you set the key on the flavor but NOT use it when you set the key on the aggregate | |
| 18:27:37 | Tengu | melwitt: yup, I have done that | |
| 18:27:45 | Tengu | but to no success until now. | |
| 18:27:46 | mriedem | and that key prefix is only used with AggregateInstanceExtraSpecsFilter | |
| 18:27:53 | cfriesen_ | was just going to mention the filter | |
| 18:27:56 | mriedem | and you have to make sure you have that enabled | |
| 18:28:08 | melwitt | Tengu: yeah, did you add that filter to your configured filters for the FilterScheduler? | |
| 18:28:12 | melwitt | in nova.conf | |
| 18:28:16 | Tengu | it's enabled. should it be in the first position? | |
| 18:28:20 | mriedem | no | |
| 18:28:24 | Tengu | melwitt: yep, it's present | |
| 18:28:28 | mriedem | order only matters for performance | |
| 18:28:43 | Tengu | and I rebooted the controllers in order to ensure all is running at the latest config version | |
| 18:28:47 | melwitt | Tengu: no but you will want to check nova-scheduler logs to make sure some other filter isn't rejecting it | |
| 18:28:48 | Tengu | mriedem: hmm ok. | |
| 18:29:12 | melwitt | at DEBUG log level. it's possible something else is going wrong and not the key match for the metadata | |
| 18:29:16 | Tengu | melwitt: yup, but I didn't see anything. the instance "directory" was created on the right node in /var/lib/nova/instances | |
| 18:29:31 | melwitt | if you're getting NoValidHost you should see something | |
| 18:29:49 | Tengu | but after a while, paff, directory is removed, and crash, "no host found"… although it actually HAD found a host | |
| 18:30:00 | melwitt | unless a compute host rejected the request in which case you should see an error in the nova-compute logs or the nova-conductor logs | |
| 18:30:08 | Tengu | hmmm. | |
| 18:30:29 | Tengu | will check that one. | |
| 18:30:53 | melwitt | the way it works is if scheduling filters all pass, it goes to nova-compute, if something fails while it builds it, it will tear it down, log stuff, and try to reschedule to another host if you have retries configured | |
| 18:31:05 | Tengu | what would be the patter of a rejection in nova-compute.log ? | |
| 18:31:16 | Tengu | hmm ok. | |
| 18:31:21 | Tengu | I have the retryfilter | |
| 18:31:23 | melwitt | should see something logged at ERROR level I think | |
| 18:31:28 | melwitt | in nova-compute | |
| 18:31:29 | Tengu | think this one will try to re-schedule | |
| 18:31:45 | Tengu | duh | |
| 18:32:05 | Tengu | corrupted image download o_O | |
| 18:32:47 | Tengu | that might explain a bit. but that would point the glance storage | |
| 18:33:17 | openstackgerrit | Eric Berglund proposed openstack/nova master: PowerVM Driver: config drive https://review.openstack.org/409404 | |
| 18:33:18 | openstackgerrit | Eric Berglund proposed openstack/nova master: WIP(5): PowerVM driver: ovs vif https://review.openstack.org/422512 | |
| 18:33:31 | Tengu | although… hmm. timestamp doesn't really match. will dig a bit more. | |
| 18:34:31 | melwitt | Tengu: yeah, so you have other issues there. but as long as the instance is always landing on the host where the aggregate meta matches the flavor, you know at least the extra specs filtering is working correctly | |
| 18:34:44 | melwitt | (for your original concern) | |
| 18:35:11 | Tengu | melwitt: right. | |
| 18:35:37 | Tengu | so my debug steps weren't that wrong. I should have had a better look to the nova-compute.log file though. | |
| 18:36:44 | mriedem | you can also trace the request id and/or instance id through the logs if you have your logs pumped to an ELK stack | |
| 18:37:05 | mriedem | or journald like in devstack | |
| 18:38:29 | Tengu | for now we don't have an ELK (it will run on the openstack… well, yes, that might cause some issues at some point ;)). | |
| 18:38:34 | Tengu | but we want to do that, yep. | |
| 18:39:01 | dansmith | mriedem: https://imgur.com/a/FY7Oq | |
| 18:39:14 | dansmith | mriedem: over about 300 runs, my patch is consistently faster than master | |
| 18:39:37 | mriedem | oh that's w/o the policy fix :) | |
| 18:39:40 | mriedem | i was like, wtf | |
| 18:39:43 | dansmith | yes | |
| 18:39:47 | Tengu | but the image corruption is the best hint for now. Have to check why - the ceph cluster isn't a cluster for now and we have some failed disks on it, so it can explain a lot. it's not in prod for now, this also explain some issues | |
| 18:41:21 | mriedem | dansmith: throw that in https://etherpad.openstack.org/p/nova-instance-list somewhere so we don't lose it | |
| 18:42:06 | stvnoyes | hi mriedem, if you get a change to re-review https://review.openstack.org/#/c/463987/ it would be great. I am on vacation next week so there's still some time this week for me to turn the review around again if it's needed. thanks. | |
| 18:43:48 | mriedem | ok | |
| 18:44:00 | mriedem | dansmith: totally unrelated, but i'm think about throwing the ceph job in the experimental queue http://tinyurl.com/ydy3jek9 | |
| 18:44:18 | stvnoyes | johnthetubaguy: pls take a look at https://review.openstack.org/#/c/506805/ when you get a chance. it's a pretty small change, and it's needed for the cinder v3 live migrate change. thanks. | |
| 18:44:23 | dansmith | melwitt: ^ | |
| 18:44:54 | melwitt | gdi | |
| 18:46:39 | melwitt | it would take me awhile to unroll what is going on with that job | |
| 18:46:52 | mriedem | this is compared to the normal dsvm tempest job http://tinyurl.com/ydemspkl | |
| 18:47:46 | melwitt | yeah, hm. so it was tracking okay until around the 16th | |
| 18:48:14 | melwitt | I'll dig into it | |
| 18:48:19 | mriedem | well, there were also spikes in the normal job then too, just not as bad | |
| 18:48:25 | mriedem | 9/23 is where it goes nuts | |
| 18:49:10 | melwitt | yeah. last we discussed it was at the last PTG and jbernard had some TODOs but I didn't know details about what they were. I thought the first thing was something to do with a job timeout being too short | |
| 18:49:32 | mriedem | he was going to start restricting the tests | |
| 18:49:47 | melwitt | restricting in what way? | |
| 18:49:56 | mriedem | L138 https://etherpad.openstack.org/p/nova-ptg-pike | |