Earlier  
Posted Nick Remark
#openstack-nova - 2018-05-02
17:23:51 dansmith but we also always said we'd wait until we had something to measure
17:23:57 dansmith so that bill may be coming due soon
17:24:10 tssurya dansmith: yea we are measuring stuff, belmiro will give a upgrade talk at the summit
17:24:24 dansmith cool
17:25:14 melwitt to the earlier question about the cells meeting, I don't have anything other than a request for review on this patch to make affinity work with multiple cells https://review.openstack.org/540258
17:26:20 openstack Launchpad bug 1767303 in OpenStack Compute (nova) "Scheduler connects to all cells DBs to gather compute nodes info" [Undecided,New] - Assigned to Surya Seetharaman (tssurya)
17:26:20 tssurya for instance another point I wanted to know was this : https://bugs.launchpad.net/nova/+bug/1767303, we filed a new bug to only go through cells that are needed now that we have disabled the host_state cache. I will put a POC once we test the performance in our envo.
17:27:33 tssurya but yea, we can also talk in the next week's meeting after we see if scheduling time improves
17:27:36 mgagne tssurya: is there a blog post or something I can read about your migration to CellsV2? Or should I wait until summit to hear it first-hand at one of the talk?
17:28:10 tssurya mgagne: we will try to put up one soon :) we are still putting out fires
17:28:33 mgagne tssurya: all fine, I'm currently sitting in a fire too =)
17:28:54 dansmith tssurya: cool, makes sense
17:28:58 tssurya mgagne: your also on cellsv2 ?
17:28:59 dansmith mgagne: how many cells do you have?
17:29:09 tssurya you are*
17:29:10 mgagne tssurya: running CellsV1 Mitaka
17:29:24 tssurya mgagne: ack
17:29:28 mgagne dansmith: more than 1 =)
17:29:47 mgagne dansmith: I'm not (yet) concerned with performance, only the migration process with multiple cells
17:38:58 mriedem tssurya: do you know if cern was/is using the CachingScheduler before upgrading to ocata+?
17:39:14 tssurya mriedem: no we weren't and aren't
17:39:22 mriedem ok
17:41:01 tssurya dansmith: not urgent, but whenever you get the time could you re-instate the +2 on https://review.openstack.org/#/c/559299/ , its the clean up on the bp which has been held up for some time.
17:42:07 dansmith tssurya: done, sorry
17:42:16 tssurya dansmith: thank you! :)
18:10:39 mriedem jroll: is oath using the CachingScheduler for baremetal scheduling?
18:16:00 jroll mriedem: no, that would be insane, we wrote our own scheduler driver of course :|
18:16:35 jroll mriedem: (it's subclassed from filter scheduler and we're likely dropping it in ocata, so carry on with the config drops and such)
18:17:52 mriedem jroll: was more about the ML thread on this wrt caching scheduler
18:18:06 mriedem and mgagne replies about performance of baremetal scheduling with the caching scheduler
18:18:55 jroll mriedem: we'll find out soon about performance on the filter scheduler in ocata. in juno we did this crazy claims thing I proposed upstream and then dropped a while back
18:19:13 jroll https://review.openstack.org/#/c/204641/
18:20:11 mgagne I guess I'm alone going at ludicrous speed with CachingScheduler =)
18:20:24 jroll heh
18:20:47 jroll with racing fixed with the resource classes stuff, I'm not sure the scheduler performance matters so much
18:21:13 mgagne it was impossible to boot like 100 baremetal at the same time with FilterScheduler, now it's possible with CachingScheduler.
18:21:14 jroll as baremetal machines take minimum 5 minutes to deploy and boot anyway
18:21:25 jroll right
18:21:44 jroll mgagne: what version are you on?
18:21:49 mriedem mitaka
18:21:56 mgagne jroll: the best and greatest Mitaka ;)
18:22:14 mriedem mgagne: do you see a fair number of reschedules?
18:22:51 jroll ah yeah. I suspect in master, a bunch of concurrent boots are just fine
18:22:55 mgagne afaik, no. we run only one scheduler as cache is done in-memory and not shared
18:23:17 mgagne we used to see a lot of rescheduling with FilterScheduler IIRC
18:23:36 mriedem well, CachingScheduler is FilterScheduler
18:23:41 mriedem except with the cache
18:23:48 mgagne jroll: remember that OOM issue you had with scheduler with ironic? ;)
18:23:58 jroll heh
18:24:06 jroll that was loading instances and such
18:24:28 mgagne scheduler taking a gazillion bytes of ram
18:24:36 mgagne jroll: yea, great stuff
18:24:38 jroll idk, I'd like to see some benchmarking on master before we make decisions about CachingScheduler being useful on master
18:25:15 mriedem i've asked for benchmarking between placement + filterscheduler and cachingscheduler since at least ocata/pike at every summit
18:25:16 mgagne jroll: being useful is not really the issue, it's more that a migration path is not ready yet for those using CachingScheduler
18:25:25 mriedem mgagne: i'm working on that CLI right now
18:25:52 jroll right, if we backport that CLI we should be all good
18:25:54 mriedem and i did remember when replying to the ML that newton and ocata computes will write allocations to placement,
18:26:03 mriedem it was only in pike when we disabled that because the filter scheduler creates the allocations now
18:26:12 mgagne mriedem: tyvm for working on it. IMO, if it's not merged, it's not done. hence my concern with removing stuff that could break my use case.
18:26:42 mriedem mgagne: sure, but also, if you're on mitaka, you have plenty of time for this to land before you're ready to roll to it
18:27:05 mriedem i.e. the cachingscheduler and ironic_host_manager removal
18:27:28 mgagne mriedem: I'm more on a defensive stance with that regard, can never be too careful
18:28:02 mriedem i understand :)
18:28:47 mgagne on principal, I can't agree with someone suggesting a new feature/fix not yet merged to replace something that is going to be removed.
18:30:59 mgagne or principle?
18:31:10 artom The latter
18:35:20 openstackgerrit Dan Smith proposed openstack/nova master: Add CellMapping.get_by_project_id() query method https://review.openstack.org/509002
18:35:22 openstackgerrit Dan Smith proposed openstack/nova master: Make get_instance_objects_sorted() be smart about cells https://review.openstack.org/509003
18:35:23 dansmith tssurya: ^
18:35:46 tssurya dansmith: ack, and thanks
18:35:56 dansmith I think we may want to consider a flag to let you choose between "always scatter to all cells" and "only hit cells for which you have instances"
18:36:21 dansmith for your case, the latter will be a big improvement, but for people that spread across cells, doing the extra mapping lookup will be a waste
18:36:31 tssurya dansmith: sounds good, yea makes sense
18:37:03 mgagne artom: :)
18:37:28 tssurya but are you talking about doing this in the listing context alone ? or a generic flag for the scatter-gather utility
18:37:55 dansmith scatter-gather can't do it generically,
18:37:59 dansmith because it only applies to instances
18:38:08 dansmith so it would just be instance listing that would honor that
18:38:43 jroll how ready for primetime is cells v2 in ocata?
18:38:57 tssurya yea I mean scatter-gather already has the selected cells querying option, what I mean was this "only hit cells for which you have instances" sounds really appealing in a lot of places
18:39:05 tssurya like quota calculating while booting
18:39:09 tssurya calculation*
18:39:19 tssurya meant*
18:39:46 dansmith tssurya: sure, it makes sense to do it in those sorts of places, I just think it has to be baked into each one that is tenant-scoped
18:39:46 tssurya so was wondering if there was a way to flag this more generically
18:39:59 tssurya dansmith: hmm, yea
18:40:22 dansmith tssurya: anywhere we can do tenant-scoped lookups should look at that flag, if that's what you mean
18:40:29 dansmith jroll: not very
18:40:41 jroll whomp whomp
18:40:44 tssurya dansmith: yes,
18:41:58 melwitt jroll: multi cell is not a thing in ocata, I'd say
18:42:16 jroll melwitt: yeah, that was my actual question, thanks
18:43:03 melwitt and, affinity scheduling doesn't work with multi cell yet, there's a bug (I have a patch up to fix it) https://review.openstack.org/540258
18:43:35 melwitt I know that's something y'all need
18:45:02 jroll tis, thanks for that
19:02:33 melwitt hm, seeing an issue where a BFV instance being deleted fails to detach the volume because attachment_id = None http://logs.openstack.org/01/565601/2/check/tempest-full/f1004d9/controller/logs/screen-n-cpu.txt.gz#_May_02_03_50_00_958050
19:03:55 cfriesen anyone know offhand what's causing this failure when running tox against newton? http://paste.openstack.org/show/720243/
19:04:50 melwitt did you do a tox -e<whatevs> -r first?
19:04:57 melwitt usually have to do that when you switch branches

Earlier   Later