Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-27
15:01:38 dansmith even on my fast piece of hardware, placement pegs a CPU
15:01:49 dansmith mriedem: not that I've heard
15:01:58 openstackgerrit OpenStack Proposal Bot proposed openstack/os-vif stable/pike: Updated from global requirements https://review.openstack.org/493146
15:02:13 cdent yeah, we had “do some performance testing” in the weekly rp update for so long that I eventually took it out from apparent lack of interest
15:02:38 mriedem ok. our public cloud guys have made tweaks to the scheduler for performance in mitaka, and lots of those tweaks i've said, "this thing in pike should resolve/replace that" but i don't have hard evidence
15:02:44 cdent It would surprise me not one tiny bit that it is not as performant as expected, because the only testing I’m aware of was done using mostly just a database, and not any of the other bits
15:02:52 cdent and since then the queries have modified quite a bit
15:02:57 cdent and we’ve added more objects
15:03:19 mriedem cdent: so i don't think you were around yesterday when we were talking about this,
15:03:32 openstackgerrit OpenStack Proposal Bot proposed openstack/python-novaclient stable/pike: Updated from global requirements https://review.openstack.org/493187
15:03:38 mriedem but i realized, after several hours, yesterday why i couldn't burst 500 (fake) guest vms on my single node devstack
15:03:42 mriedem and they were all going novalidhost
15:03:43 cdent I had an afternoon in an attorney’s office ...
15:03:49 dansmith also keep in mind that a little slower scheduler performance compares very favorably to 10% retries in the background because we choose bad computes
15:04:04 cdent mriedem: what was the cause?
15:04:06 mriedem dansmith: that's why i'd want to compare ocata to pike
15:04:20 mriedem cdent: the ultimate cause was a 409 response from placement when putting the allocations during scheduling
15:04:24 mriedem we retry that 3 times,
15:04:30 dansmith mriedem: yeah, I'm just saying you have to consider "time to all active" not just "time to building" or something
15:04:31 mriedem but it wasn't enough apparently
15:04:36 mriedem dansmith: agree
15:04:46 cdent the 409 was for generation mismatch?
15:04:56 mriedem dansmith: with all the spinning plates, i've been thinking about starting a perf scenario etherpad, polish that up and then send out to see if people can help
15:05:11 dansmith cdent: it's allocation, so it's the internal rp generation conflict I think
15:05:15 mriedem cdent: this is the response, "There was a conflict when trying to complete your request.\n\n Inventory changed while attempting to allocate: Another thread concurrently updated the data. Please retry your update"
15:05:34 dansmith cdent: and I'm not sure why placement isn't just retrying that for us
15:05:38 cdent https://github.com/openstack/nova/blob/master/nova/objects/resource_provider.py#L1837-L1839
15:05:53 mriedem i didn't see any actual inventory updates from the virt driver, which shouldn't happen since the inventory in this case is static
15:05:53 dansmith ah sweet
15:05:58 mriedem so the message was a bit misleading
15:05:58 cdent dansmith: yeah, that’s what I meant by generation mismatch
15:06:13 dansmith cdent: yeah, I know where it's happening, but hadn't seen that TODO from jaypipes
15:06:15 dansmith so that's cool
15:06:23 dansmith cdent: I'm not sure why we're hitting it though,
15:06:35 dansmith cdent: since it's a single thread of allocating for things, with no inventory updates coming from the compute
15:06:44 dansmith so I'm not sure what is racing
15:07:03 cdent thinking out loud: every time we write an allocation we update the generation
15:07:10 dansmith right
15:07:14 cdent and we compare the generation with what the generation was before we entered the transaction
15:07:23 cdent so we race to get the transaction
15:07:24 dansmith and we conflict if something else changes the generation while we're trying to
15:07:33 gibi mriedem: I don't think we saw a real race on master see my comment in the bughttps://bugs.launchpad.net/nova/+bug/1719915/comments/1
15:07:36 openstack Launchpad bug 1719915 in OpenStack Compute (nova) "test_live_migrate_delete race fail when checking allocations: MismatchError: 2 != 1" [Medium,Confirmed] - Assigned to Balazs Gibizer (balazs-gibizer)
15:08:04 cdent we create an rp object for each allocation at the http layer
15:08:10 cdent that’s the generation that’s being used
15:09:03 mriedem gibi: http://logstash.openstack.org/#dashboard/file/logstash.json?query=message%3A%5C%22%5Bnova.api.openstack.requestlog%5D%20127.0.0.1%20%5C%5C%5C%22DELETE%20%2Fv2.1%2F%5C%22%20AND%20message%3A%5C%22%2Fmigrations%2F1%5C%22%20AND%20tags%3A%5C%22console%5C%22&from=7d
15:09:24 cdent yeah
15:10:06 cdent most straightforward thing to do, presumably, is to do the TODO, and retry 10 times server side, so the client would be effectively retrying 30 times?
15:10:41 dansmith cdent: sure, we should be retrying server side,
15:10:57 dansmith cdent: my point is I don't know why we'd be hitting this need to retry with a single thread of allocations
15:11:15 cdent (efried I haven’t got an opinion on that conf/utils.py issue)
15:11:16 mriedem right, we process the instances in a for loop in the scheduler
15:11:30 efried cdent Ack, thanks for looking.
15:11:31 mriedem so we're put'ing the allocations to the same host, but in order
15:11:43 mriedem and the compute shouldn't be changing any inventory since it's static
15:11:59 mriedem i grep'ed the logs for PUT.*inventories and there was nothing
15:12:11 cdent is there anything else putting allocations?
15:12:24 dansmith cdent: no, single 100-instance boot, so one for loop
15:12:25 mriedem would have to audit that, i didn't dig yet
15:12:35 mriedem cdent: like the compute?
15:12:38 dansmith I mean.. "shouldn't be"
15:12:44 mriedem right, nothing else shoudl be
15:12:49 mriedem since we're not doing any moves or anything
15:12:50 cdent yeah, I’m wondering if we left something else somewhere that we forgot about?
15:13:01 cdent I know it’s not supposed to be, but given everything...
15:13:03 dansmith mriedem: remember I suggested to see if the compute was doing ocata fallback behavior for some reason
15:13:08 gibi mriedem: OK, thats a different failure than the what originally was pasted to the bug report. I continue digging...
15:13:19 cdent another possibility is that uwsgi is (somehow, who knows) letting things get out of order
15:13:29 mriedem dansmith: do we log anything specific in that case?
15:13:40 mriedem i see a buttload of the "we're on a pike compute with all pike computes, so not healing allocations" all the time
15:13:41 dansmith mriedem: placement will log it
15:13:51 dansmith mriedem: okay then that probably means it's not
15:13:59 mriedem but ^ is from the periodic
15:14:18 mriedem there are paths in the RT that go into that code w/o consciously passing the has_ocata_computes flag
15:14:22 mriedem but i think it defaults to False anyway
15:15:13 cdent mriedem: have you got a set of logs you can make available?
15:15:18 openstackgerrit OpenStack Proposal Bot proposed openstack/os-vif stable/pike: Updated from global requirements https://review.openstack.org/493146
15:16:02 cdent this doesn’t feel like something it’s going to be easy to reason about without some files to grep
15:16:48 dansmith cdent: it should be pretty easy to reproduce (or not) in a devstack and then you can instrument the code as needed
15:16:58 openstackgerrit OpenStack Proposal Bot proposed openstack/python-novaclient stable/pike: Updated from global requirements https://review.openstack.org/493187
15:17:12 mriedem cdent: no, it's all local
15:17:19 mriedem well, in this devstack vm which is not local
15:17:28 mriedem but yeah i have the local.conf for the devstack if you want to reproduce
15:17:32 cdent mriedem: sure, but you have tar and such?
15:17:41 mriedem yeah
15:18:16 mriedem is there a standard way to tar up the journald logs?
15:18:33 cdent balls, I forgot about journald, meh
15:18:46 mriedem it's tar'ed up in devstack-gate
15:18:52 mriedem so i can just copy whatever we do in CI
15:19:46 openstackgerrit Rodolfo Alonso Hernandez proposed openstack/nova master: Change 'InstancePCIRequest' spec field https://review.openstack.org/449257
15:19:47 cdent I can’t really look with any real attention until about 3 hours from now, but if you get a chance to do it, that’s great it will useful, if not, just the local.conf will do
15:19:57 mriedem https://github.com/openstack-infra/devstack-gate/blob/master/functions.sh#L698-L724
15:21:25 mriedem you know, i could just do this with a devstack patch
15:21:27 mriedem that's easier
15:21:34 mriedem let the ci do the work
15:25:19 mriedem needless to say, i'm doing a terrible job of reviewing code or specs, or writing specs
15:26:01 mriedem sdague: can you get this stable/pike novaclient bug fix backport? https://review.openstack.org/#/c/495901/
15:26:12 mriedem pretty nasty and we need to release it
15:26:18 sdague mriedem: looking
15:26:55 sdague +A

Earlier   Later