Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-27
15:06:15 dansmith so that's cool
15:06:23 dansmith cdent: I'm not sure why we're hitting it though,
15:06:35 dansmith cdent: since it's a single thread of allocating for things, with no inventory updates coming from the compute
15:06:44 dansmith so I'm not sure what is racing
15:07:03 cdent thinking out loud: every time we write an allocation we update the generation
15:07:10 dansmith right
15:07:14 cdent and we compare the generation with what the generation was before we entered the transaction
15:07:23 cdent so we race to get the transaction
15:07:24 dansmith and we conflict if something else changes the generation while we're trying to
15:07:33 gibi mriedem: I don't think we saw a real race on master see my comment in the bughttps://bugs.launchpad.net/nova/+bug/1719915/comments/1
15:07:36 openstack Launchpad bug 1719915 in OpenStack Compute (nova) "test_live_migrate_delete race fail when checking allocations: MismatchError: 2 != 1" [Medium,Confirmed] - Assigned to Balazs Gibizer (balazs-gibizer)
15:08:04 cdent we create an rp object for each allocation at the http layer
15:08:10 cdent that’s the generation that’s being used
15:09:03 mriedem gibi: http://logstash.openstack.org/#dashboard/file/logstash.json?query=message%3A%5C%22%5Bnova.api.openstack.requestlog%5D%20127.0.0.1%20%5C%5C%5C%22DELETE%20%2Fv2.1%2F%5C%22%20AND%20message%3A%5C%22%2Fmigrations%2F1%5C%22%20AND%20tags%3A%5C%22console%5C%22&from=7d
15:09:24 cdent yeah
15:10:06 cdent most straightforward thing to do, presumably, is to do the TODO, and retry 10 times server side, so the client would be effectively retrying 30 times?
15:10:41 dansmith cdent: sure, we should be retrying server side,
15:10:57 dansmith cdent: my point is I don't know why we'd be hitting this need to retry with a single thread of allocations
15:11:15 cdent (efried I haven’t got an opinion on that conf/utils.py issue)
15:11:16 mriedem right, we process the instances in a for loop in the scheduler
15:11:30 efried cdent Ack, thanks for looking.
15:11:31 mriedem so we're put'ing the allocations to the same host, but in order
15:11:43 mriedem and the compute shouldn't be changing any inventory since it's static
15:11:59 mriedem i grep'ed the logs for PUT.*inventories and there was nothing
15:12:11 cdent is there anything else putting allocations?
15:12:24 dansmith cdent: no, single 100-instance boot, so one for loop
15:12:25 mriedem would have to audit that, i didn't dig yet
15:12:35 mriedem cdent: like the compute?
15:12:38 dansmith I mean.. "shouldn't be"
15:12:44 mriedem right, nothing else shoudl be
15:12:49 mriedem since we're not doing any moves or anything
15:12:50 cdent yeah, I’m wondering if we left something else somewhere that we forgot about?
15:13:01 cdent I know it’s not supposed to be, but given everything...
15:13:03 dansmith mriedem: remember I suggested to see if the compute was doing ocata fallback behavior for some reason
15:13:08 gibi mriedem: OK, thats a different failure than the what originally was pasted to the bug report. I continue digging...
15:13:19 cdent another possibility is that uwsgi is (somehow, who knows) letting things get out of order
15:13:29 mriedem dansmith: do we log anything specific in that case?
15:13:40 mriedem i see a buttload of the "we're on a pike compute with all pike computes, so not healing allocations" all the time
15:13:41 dansmith mriedem: placement will log it
15:13:51 dansmith mriedem: okay then that probably means it's not
15:13:59 mriedem but ^ is from the periodic
15:14:18 mriedem there are paths in the RT that go into that code w/o consciously passing the has_ocata_computes flag
15:14:22 mriedem but i think it defaults to False anyway
15:15:13 cdent mriedem: have you got a set of logs you can make available?
15:15:18 openstackgerrit OpenStack Proposal Bot proposed openstack/os-vif stable/pike: Updated from global requirements https://review.openstack.org/493146
15:16:02 cdent this doesn’t feel like something it’s going to be easy to reason about without some files to grep
15:16:48 dansmith cdent: it should be pretty easy to reproduce (or not) in a devstack and then you can instrument the code as needed
15:16:58 openstackgerrit OpenStack Proposal Bot proposed openstack/python-novaclient stable/pike: Updated from global requirements https://review.openstack.org/493187
15:17:12 mriedem cdent: no, it's all local
15:17:19 mriedem well, in this devstack vm which is not local
15:17:28 mriedem but yeah i have the local.conf for the devstack if you want to reproduce
15:17:32 cdent mriedem: sure, but you have tar and such?
15:17:41 mriedem yeah
15:18:16 mriedem is there a standard way to tar up the journald logs?
15:18:33 cdent balls, I forgot about journald, meh
15:18:46 mriedem it's tar'ed up in devstack-gate
15:18:52 mriedem so i can just copy whatever we do in CI
15:19:46 openstackgerrit Rodolfo Alonso Hernandez proposed openstack/nova master: Change 'InstancePCIRequest' spec field https://review.openstack.org/449257
15:19:47 cdent I can’t really look with any real attention until about 3 hours from now, but if you get a chance to do it, that’s great it will useful, if not, just the local.conf will do
15:19:57 mriedem https://github.com/openstack-infra/devstack-gate/blob/master/functions.sh#L698-L724
15:21:25 mriedem you know, i could just do this with a devstack patch
15:21:27 mriedem that's easier
15:21:34 mriedem let the ci do the work
15:25:19 mriedem needless to say, i'm doing a terrible job of reviewing code or specs, or writing specs
15:26:01 mriedem sdague: can you get this stable/pike novaclient bug fix backport? https://review.openstack.org/#/c/495901/
15:26:12 mriedem pretty nasty and we need to release it
15:26:18 sdague mriedem: looking
15:26:55 sdague +A
15:27:32 dansmith mriedem: speaking of reviewing, I'm not sure what else to do on the base switchover patch since the difference is not measurable on my box. I'm poking at the fault thing in the later patch, but we can just strip that out and work on it in parallel
15:27:52 mriedem i haven't reviewed the actual change yet
15:28:04 mriedem was just getting 1000 active vms locally to test the fault thing we talked about last night
15:28:14 mriedem i agree that what we found last night, for numbers, isn't worth holding things up
15:28:28 mriedem sdague: thanks
15:28:34 dansmith mriedem: okay, like I said in the etherpad, I don't think we're doing the fault thing in that patch yet
15:28:47 dansmith and applying the one that does it definitely has an impact
15:29:15 mriedem negative impact?
15:29:37 dansmith yeah
15:30:05 mriedem because we've added a new unconditional join i suppose
15:30:19 dansmith it's not a new join, but it's a new query yeah
15:30:25 mriedem yeah, was just thinking that
15:30:43 dansmith I'm messing with the later patch though to see if I can squash it out
15:30:45 mriedem like you said, we could plumb that in the db api
15:30:53 openstackgerrit OpenStack Proposal Bot proposed openstack/os-vif stable/newton: Updated from global requirements https://review.openstack.org/373293
15:31:05 dansmith that's what I'm doing yeah
15:31:05 mriedem if faults is in expected_attrs, get the instances in deleted/error and include the faults on those?
15:32:04 dansmith yes
15:36:47 cdent do dansmith and mriedem have an exciting new etherpad they are willing to share?
15:37:06 dansmith cdent: https://etherpad.openstack.org/p/nova-instance-list
15:37:12 dansmith cdent: but that's just incidental to him finding the problem
15:37:21 dansmith it's not about the thing we were describing
15:37:35 cdent you know I’m a total junkie for all the info
15:37:55 mriedem started as getting a benchmark for testing before and after the big instance list change
15:38:10 mriedem but have also added todos for weird things we've seen, like the concurrent update detected thing with 500 instances
15:42:09 mriedem sdague: can i put [[post-config|$NOVA_CONF]] in stackrc?
15:42:44 cdent mriedem, dansmith, (and sean mooney): If any of you get a chance to look at the discussion on https://review.openstack.org/#/c/504540/ (limiting GET /allocation_candidates ) it’s gotten to the point where we are trying to decide what it is that we are actually optimizing for, so could do with more input
15:43:17 sdague mriedem: in local.conf
15:43:42 mriedem sdague: yeah, locally, but this is for running something through ci
15:44:51 sdague there isn't really anyway to jam it into stackrc, you need to do project-config changes for stuff like that
15:45:15 sdague if this is just for hacktastic stuff, just iniset whatever you want in lib/nova
15:45:16 mriedem ok, just checking, i can hack this other ways

Earlier   Later