Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-27
18:17:43 melwitt Tengu: I think if you tag your flavors with extra_specs and then use the AggregateInstanceExtraSpecsFilter you can do what you want
18:18:12 Tengu after that, the flavor were created, and a metadata was added in the form "gen1=true" for m1.medium, and "gen2=true" for m2.medium
18:18:18 mriedem the problem is,
18:18:31 mriedem flavor1 is associated to agg1 and flavor2 is associated to agg2,
18:18:36 mriedem but that doesn't exclude agg2 from using flavor1
18:18:38 mriedem and vice versa
18:18:42 mriedem that's the strict isolation problme
18:18:42 Tengu hmm ok.
18:18:44 mriedem *problem
18:18:56 Tengu not a really big issue - for now, we have "no host found" in fact
18:19:01 melwitt I thought if the flavors were tagged it would require that key to pass?
18:19:24 mriedem honestly i'd have to re-read https://review.openstack.org/#/c/381912/
18:19:33 melwitt I'm reading it again now
18:19:35 mriedem i am definitely not an expert here on the existing capabilities and gaps
18:19:36 Tengu melwitt: same for me - actually, for now, we're unable to start any instance because it doesn't find any host to run it
18:20:45 Tengu mriedem: but maybe it's "just" the metadata format that fails me. is there any doc for that?
18:20:57 melwitt Tengu: and you added gen1=true and gen2=true to your host aggregates?
18:21:11 Tengu yup, as a metadata as well
18:21:40 mriedem the now deleted ops guide might have had something specific for this
18:21:48 Tengu :'(
18:22:14 cdent it got moved to the wiki?
18:22:14 Tengu I found a doc saying the metadata on the flavor should be in the form aggregate_instance_extra_specs:gen1='true'
18:22:20 Tengu but that doesn't work either
18:22:20 cfriesen_ mriedem: dansmith: just saw the mention of microversion 2.47...there was already a call to "instance.get_flavor()" previously, so I had assumed it would get the whole flavor. I suspect you're right that it's lazy-loading extra-specs.
18:22:33 mriedem it would be in here if it existed https://docs.openstack.org/nova/latest/admin/index.html
18:22:40 dansmith cfriesen_: no it's policy
18:22:55 dansmith cfriesen_: the policy is checked per instance now, which is an fs call at least
18:23:09 cfriesen_ dansmith: ah...I had a networking glitch, missed some irc.
18:23:43 cfriesen_ dansmith: fix is what, cache the policy?
18:23:56 dansmith cfriesen_: check once per list and not once per instance
18:24:01 dansmith cfriesen_: I'm cooking it up now
18:24:03 dansmith smells like bacon
18:24:16 cfriesen_ dansmith: do we even need that? couldn't we check it the first time and cache it?
18:24:26 dansmith cfriesen_: that's what I just said
18:24:37 cfriesen_ I meant the first time on process startup
18:24:54 mriedem Tengu: i've seen a better doc than that red hat one, sec
18:25:12 dansmith cfriesen_: it depends per request
18:25:46 Tengu mriedem: that would be nice :)
18:25:47 cfriesen_ dansmith: ah, of course
18:25:53 melwitt Tengu: I found this doc https://docs.openstack.org/ocata/config-reference/compute/schedulers.html#host-aggregates
18:26:04 dansmith mriedem: confirmed the knee in the same place on master
18:26:05 Tengu I've also followed https://blog.russellbryant.net/2013/05/21/availability-zones-and-host-aggregates-in-openstack-compute-nova/ - but failed.
18:26:27 Tengu melwitt: ah, ocata, might work, pike is just one version ahead. will check that, thanks!
18:26:43 mriedem melwitt: yeah https://docs.openstack.org/ocata/config-reference/compute/schedulers.html#example-specify-compute-hosts-with-ssds
18:26:57 mriedem openstack flavor set --property aggregate_instance_extra_specs:ssd=true ssd.large
18:27:08 Tengu duh… ok, I was also on that one -.-'
18:27:17 melwitt Tengu: the main thing I saw ppl run into a snag is that you apparently have to use that prefix when you set the key on the flavor but NOT use it when you set the key on the aggregate
18:27:37 Tengu melwitt: yup, I have done that
18:27:45 Tengu but to no success until now.
18:27:46 mriedem and that key prefix is only used with AggregateInstanceExtraSpecsFilter
18:27:53 cfriesen_ was just going to mention the filter
18:27:56 mriedem and you have to make sure you have that enabled
18:28:08 melwitt Tengu: yeah, did you add that filter to your configured filters for the FilterScheduler?
18:28:12 melwitt in nova.conf
18:28:16 Tengu it's enabled. should it be in the first position?
18:28:20 mriedem no
18:28:24 Tengu melwitt: yep, it's present
18:28:28 mriedem order only matters for performance
18:28:43 Tengu and I rebooted the controllers in order to ensure all is running at the latest config version
18:28:47 melwitt Tengu: no but you will want to check nova-scheduler logs to make sure some other filter isn't rejecting it
18:28:48 Tengu mriedem: hmm ok.
18:29:12 melwitt at DEBUG log level. it's possible something else is going wrong and not the key match for the metadata
18:29:16 Tengu melwitt: yup, but I didn't see anything. the instance "directory" was created on the right node in /var/lib/nova/instances
18:29:31 melwitt if you're getting NoValidHost you should see something
18:29:49 Tengu but after a while, paff, directory is removed, and crash, "no host found"… although it actually HAD found a host
18:30:00 melwitt unless a compute host rejected the request in which case you should see an error in the nova-compute logs or the nova-conductor logs
18:30:08 Tengu hmmm.
18:30:29 Tengu will check that one.
18:30:53 melwitt the way it works is if scheduling filters all pass, it goes to nova-compute, if something fails while it builds it, it will tear it down, log stuff, and try to reschedule to another host if you have retries configured
18:31:05 Tengu what would be the patter of a rejection in nova-compute.log ?
18:31:16 Tengu hmm ok.
18:31:21 Tengu I have the retryfilter
18:31:23 melwitt should see something logged at ERROR level I think
18:31:28 melwitt in nova-compute
18:31:29 Tengu think this one will try to re-schedule
18:31:45 Tengu duh
18:32:05 Tengu corrupted image download o_O
18:32:47 Tengu that might explain a bit. but that would point the glance storage
18:33:17 openstackgerrit Eric Berglund proposed openstack/nova master: PowerVM Driver: config drive https://review.openstack.org/409404
18:33:18 openstackgerrit Eric Berglund proposed openstack/nova master: WIP(5): PowerVM driver: ovs vif https://review.openstack.org/422512
18:33:31 Tengu although… hmm. timestamp doesn't really match. will dig a bit more.
18:34:31 melwitt Tengu: yeah, so you have other issues there. but as long as the instance is always landing on the host where the aggregate meta matches the flavor, you know at least the extra specs filtering is working correctly
18:34:44 melwitt (for your original concern)
18:35:11 Tengu melwitt: right.
18:35:37 Tengu so my debug steps weren't that wrong. I should have had a better look to the nova-compute.log file though.
18:36:44 mriedem you can also trace the request id and/or instance id through the logs if you have your logs pumped to an ELK stack
18:37:05 mriedem or journald like in devstack
18:38:29 Tengu for now we don't have an ELK (it will run on the openstack… well, yes, that might cause some issues at some point ;)).
18:38:34 Tengu but we want to do that, yep.
18:39:01 dansmith mriedem: https://imgur.com/a/FY7Oq
18:39:14 dansmith mriedem: over about 300 runs, my patch is consistently faster than master
18:39:37 mriedem oh that's w/o the policy fix :)
18:39:40 mriedem i was like, wtf
18:39:43 dansmith yes
18:39:47 Tengu but the image corruption is the best hint for now. Have to check why - the ceph cluster isn't a cluster for now and we have some failed disks on it, so it can explain a lot. it's not in prod for now, this also explain some issues
18:41:21 mriedem dansmith: throw that in https://etherpad.openstack.org/p/nova-instance-list somewhere so we don't lose it
18:42:06 stvnoyes hi mriedem, if you get a change to re-review https://review.openstack.org/#/c/463987/ it would be great. I am on vacation next week so there's still some time this week for me to turn the review around again if it's needed. thanks.
18:43:48 mriedem ok
18:44:00 mriedem dansmith: totally unrelated, but i'm think about throwing the ceph job in the experimental queue http://tinyurl.com/ydy3jek9
18:44:18 stvnoyes johnthetubaguy: pls take a look at https://review.openstack.org/#/c/506805/ when you get a chance. it's a pretty small change, and it's needed for the cinder v3 live migrate change. thanks.

Earlier   Later