| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-27 | |||
| 18:12:46 | mriedem | well, read the spec first and confirm if that's what you're asking for | |
| 18:13:49 | Tengu | looks like what we want, yes. but that's strange, I found some doc, even at Redhat, saying "it works" but without proper example. | |
| 18:14:04 | mriedem | to summarize, we talked about this at the pike ptg in february, we needed to have the various use cases documented in the spec to make sure the solution would cover them, and there were at least 2 stakeholders in the room saying, "we have an out of tree filter that does something like this" and we said, ok read this and tell us if it will replace your out of tree filter, and those people never replied to ack that it does | |
| 18:14:35 | Tengu | erf | |
| 18:14:46 | Tengu | may I explain what I did? | |
| 18:15:13 | Tengu | and point to the doc I followed - maybe a solution might be found | |
| 18:16:25 | Tengu | mriedem: I followed https://access.redhat.com/documentation/en-US/Red_Hat_Enterprise_Linux_OpenStack_Platform/6/html/Administration_Guide/section-host-aggregates.html - I think that one has some equivalent in openstack "open" doc | |
| 18:17:16 | Tengu | mriedem: I activated AggregateInstanceExtraSpecsFilter filter in nova.conf, and created two aggregate - all hosts are in those aggregates (in fact, for now, only two hosts - hence once per group). | |
| 18:17:36 | Tengu | mriedem: the metadata is like "gen1=true" for first aggregate, "gen2=true" for the second. | |
| 18:17:43 | melwitt | Tengu: I think if you tag your flavors with extra_specs and then use the AggregateInstanceExtraSpecsFilter you can do what you want | |
| 18:18:12 | Tengu | after that, the flavor were created, and a metadata was added in the form "gen1=true" for m1.medium, and "gen2=true" for m2.medium | |
| 18:18:18 | mriedem | the problem is, | |
| 18:18:31 | mriedem | flavor1 is associated to agg1 and flavor2 is associated to agg2, | |
| 18:18:36 | mriedem | but that doesn't exclude agg2 from using flavor1 | |
| 18:18:38 | mriedem | and vice versa | |
| 18:18:42 | mriedem | that's the strict isolation problme | |
| 18:18:42 | Tengu | hmm ok. | |
| 18:18:44 | mriedem | *problem | |
| 18:18:56 | Tengu | not a really big issue - for now, we have "no host found" in fact | |
| 18:19:01 | melwitt | I thought if the flavors were tagged it would require that key to pass? | |
| 18:19:24 | mriedem | honestly i'd have to re-read https://review.openstack.org/#/c/381912/ | |
| 18:19:33 | melwitt | I'm reading it again now | |
| 18:19:35 | mriedem | i am definitely not an expert here on the existing capabilities and gaps | |
| 18:19:36 | Tengu | melwitt: same for me - actually, for now, we're unable to start any instance because it doesn't find any host to run it | |
| 18:20:45 | Tengu | mriedem: but maybe it's "just" the metadata format that fails me. is there any doc for that? | |
| 18:20:57 | melwitt | Tengu: and you added gen1=true and gen2=true to your host aggregates? | |
| 18:21:11 | Tengu | yup, as a metadata as well | |
| 18:21:40 | mriedem | the now deleted ops guide might have had something specific for this | |
| 18:21:48 | Tengu | :'( | |
| 18:22:14 | cdent | it got moved to the wiki? | |
| 18:22:14 | Tengu | I found a doc saying the metadata on the flavor should be in the form aggregate_instance_extra_specs:gen1='true' | |
| 18:22:20 | Tengu | but that doesn't work either | |
| 18:22:20 | cfriesen_ | mriedem: dansmith: just saw the mention of microversion 2.47...there was already a call to "instance.get_flavor()" previously, so I had assumed it would get the whole flavor. I suspect you're right that it's lazy-loading extra-specs. | |
| 18:22:33 | mriedem | it would be in here if it existed https://docs.openstack.org/nova/latest/admin/index.html | |
| 18:22:40 | dansmith | cfriesen_: no it's policy | |
| 18:22:55 | dansmith | cfriesen_: the policy is checked per instance now, which is an fs call at least | |
| 18:23:09 | cfriesen_ | dansmith: ah...I had a networking glitch, missed some irc. | |
| 18:23:43 | cfriesen_ | dansmith: fix is what, cache the policy? | |
| 18:23:56 | dansmith | cfriesen_: check once per list and not once per instance | |
| 18:24:01 | dansmith | cfriesen_: I'm cooking it up now | |
| 18:24:03 | dansmith | smells like bacon | |
| 18:24:16 | cfriesen_ | dansmith: do we even need that? couldn't we check it the first time and cache it? | |
| 18:24:26 | dansmith | cfriesen_: that's what I just said | |
| 18:24:37 | cfriesen_ | I meant the first time on process startup | |
| 18:24:54 | mriedem | Tengu: i've seen a better doc than that red hat one, sec | |
| 18:25:12 | dansmith | cfriesen_: it depends per request | |
| 18:25:46 | Tengu | mriedem: that would be nice :) | |
| 18:25:47 | cfriesen_ | dansmith: ah, of course | |
| 18:25:53 | melwitt | Tengu: I found this doc https://docs.openstack.org/ocata/config-reference/compute/schedulers.html#host-aggregates | |
| 18:26:04 | dansmith | mriedem: confirmed the knee in the same place on master | |
| 18:26:05 | Tengu | I've also followed https://blog.russellbryant.net/2013/05/21/availability-zones-and-host-aggregates-in-openstack-compute-nova/ - but failed. | |
| 18:26:27 | Tengu | melwitt: ah, ocata, might work, pike is just one version ahead. will check that, thanks! | |
| 18:26:43 | mriedem | melwitt: yeah https://docs.openstack.org/ocata/config-reference/compute/schedulers.html#example-specify-compute-hosts-with-ssds | |
| 18:26:57 | mriedem | openstack flavor set --property aggregate_instance_extra_specs:ssd=true ssd.large | |
| 18:27:08 | Tengu | duh… ok, I was also on that one -.-' | |
| 18:27:17 | melwitt | Tengu: the main thing I saw ppl run into a snag is that you apparently have to use that prefix when you set the key on the flavor but NOT use it when you set the key on the aggregate | |
| 18:27:37 | Tengu | melwitt: yup, I have done that | |
| 18:27:45 | Tengu | but to no success until now. | |
| 18:27:46 | mriedem | and that key prefix is only used with AggregateInstanceExtraSpecsFilter | |
| 18:27:53 | cfriesen_ | was just going to mention the filter | |
| 18:27:56 | mriedem | and you have to make sure you have that enabled | |
| 18:28:08 | melwitt | Tengu: yeah, did you add that filter to your configured filters for the FilterScheduler? | |
| 18:28:12 | melwitt | in nova.conf | |
| 18:28:16 | Tengu | it's enabled. should it be in the first position? | |
| 18:28:20 | mriedem | no | |
| 18:28:24 | Tengu | melwitt: yep, it's present | |
| 18:28:28 | mriedem | order only matters for performance | |
| 18:28:43 | Tengu | and I rebooted the controllers in order to ensure all is running at the latest config version | |
| 18:28:47 | melwitt | Tengu: no but you will want to check nova-scheduler logs to make sure some other filter isn't rejecting it | |
| 18:28:48 | Tengu | mriedem: hmm ok. | |
| 18:29:12 | melwitt | at DEBUG log level. it's possible something else is going wrong and not the key match for the metadata | |
| 18:29:16 | Tengu | melwitt: yup, but I didn't see anything. the instance "directory" was created on the right node in /var/lib/nova/instances | |
| 18:29:31 | melwitt | if you're getting NoValidHost you should see something | |
| 18:29:49 | Tengu | but after a while, paff, directory is removed, and crash, "no host found"… although it actually HAD found a host | |
| 18:30:00 | melwitt | unless a compute host rejected the request in which case you should see an error in the nova-compute logs or the nova-conductor logs | |
| 18:30:08 | Tengu | hmmm. | |
| 18:30:29 | Tengu | will check that one. | |
| 18:30:53 | melwitt | the way it works is if scheduling filters all pass, it goes to nova-compute, if something fails while it builds it, it will tear it down, log stuff, and try to reschedule to another host if you have retries configured | |
| 18:31:05 | Tengu | what would be the patter of a rejection in nova-compute.log ? | |
| 18:31:16 | Tengu | hmm ok. | |
| 18:31:21 | Tengu | I have the retryfilter | |
| 18:31:23 | melwitt | should see something logged at ERROR level I think | |
| 18:31:28 | melwitt | in nova-compute | |
| 18:31:29 | Tengu | think this one will try to re-schedule | |
| 18:31:45 | Tengu | duh | |
| 18:32:05 | Tengu | corrupted image download o_O | |
| 18:32:47 | Tengu | that might explain a bit. but that would point the glance storage | |
| 18:33:17 | openstackgerrit | Eric Berglund proposed openstack/nova master: PowerVM Driver: config drive https://review.openstack.org/409404 | |
| 18:33:18 | openstackgerrit | Eric Berglund proposed openstack/nova master: WIP(5): PowerVM driver: ovs vif https://review.openstack.org/422512 | |
| 18:33:31 | Tengu | although… hmm. timestamp doesn't really match. will dig a bit more. | |
| 18:34:31 | melwitt | Tengu: yeah, so you have other issues there. but as long as the instance is always landing on the host where the aggregate meta matches the flavor, you know at least the extra specs filtering is working correctly | |
| 18:34:44 | melwitt | (for your original concern) | |
| 18:35:11 | Tengu | melwitt: right. | |
| 18:35:37 | Tengu | so my debug steps weren't that wrong. I should have had a better look to the nova-compute.log file though. | |
| 18:36:44 | mriedem | you can also trace the request id and/or instance id through the logs if you have your logs pumped to an ELK stack | |
| 18:37:05 | mriedem | or journald like in devstack | |
| 18:38:29 | Tengu | for now we don't have an ELK (it will run on the openstack… well, yes, that might cause some issues at some point ;)). | |
| 18:38:34 | Tengu | but we want to do that, yep. | |
| 18:39:01 | dansmith | mriedem: https://imgur.com/a/FY7Oq | |
| 18:39:14 | dansmith | mriedem: over about 300 runs, my patch is consistently faster than master | |