| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-27 | |||
| 15:07:03 | cdent | thinking out loud: every time we write an allocation we update the generation | |
| 15:07:10 | dansmith | right | |
| 15:07:14 | cdent | and we compare the generation with what the generation was before we entered the transaction | |
| 15:07:23 | cdent | so we race to get the transaction | |
| 15:07:24 | dansmith | and we conflict if something else changes the generation while we're trying to | |
| 15:07:33 | gibi | mriedem: I don't think we saw a real race on master see my comment in the bughttps://bugs.launchpad.net/nova/+bug/1719915/comments/1 | |
| 15:07:36 | openstack | Launchpad bug 1719915 in OpenStack Compute (nova) "test_live_migrate_delete race fail when checking allocations: MismatchError: 2 != 1" [Medium,Confirmed] - Assigned to Balazs Gibizer (balazs-gibizer) | |
| 15:08:04 | cdent | we create an rp object for each allocation at the http layer | |
| 15:08:10 | cdent | that’s the generation that’s being used | |
| 15:09:03 | mriedem | gibi: http://logstash.openstack.org/#dashboard/file/logstash.json?query=message%3A%5C%22%5Bnova.api.openstack.requestlog%5D%20127.0.0.1%20%5C%5C%5C%22DELETE%20%2Fv2.1%2F%5C%22%20AND%20message%3A%5C%22%2Fmigrations%2F1%5C%22%20AND%20tags%3A%5C%22console%5C%22&from=7d | |
| 15:09:24 | cdent | yeah | |
| 15:10:06 | cdent | most straightforward thing to do, presumably, is to do the TODO, and retry 10 times server side, so the client would be effectively retrying 30 times? | |
| 15:10:41 | dansmith | cdent: sure, we should be retrying server side, | |
| 15:10:57 | dansmith | cdent: my point is I don't know why we'd be hitting this need to retry with a single thread of allocations | |
| 15:11:15 | cdent | (efried I haven’t got an opinion on that conf/utils.py issue) | |
| 15:11:16 | mriedem | right, we process the instances in a for loop in the scheduler | |
| 15:11:30 | efried | cdent Ack, thanks for looking. | |
| 15:11:31 | mriedem | so we're put'ing the allocations to the same host, but in order | |
| 15:11:43 | mriedem | and the compute shouldn't be changing any inventory since it's static | |
| 15:11:59 | mriedem | i grep'ed the logs for PUT.*inventories and there was nothing | |
| 15:12:11 | cdent | is there anything else putting allocations? | |
| 15:12:24 | dansmith | cdent: no, single 100-instance boot, so one for loop | |
| 15:12:25 | mriedem | would have to audit that, i didn't dig yet | |
| 15:12:35 | mriedem | cdent: like the compute? | |
| 15:12:38 | dansmith | I mean.. "shouldn't be" | |
| 15:12:44 | mriedem | right, nothing else shoudl be | |
| 15:12:49 | mriedem | since we're not doing any moves or anything | |
| 15:12:50 | cdent | yeah, I’m wondering if we left something else somewhere that we forgot about? | |
| 15:13:01 | cdent | I know it’s not supposed to be, but given everything... | |
| 15:13:03 | dansmith | mriedem: remember I suggested to see if the compute was doing ocata fallback behavior for some reason | |
| 15:13:08 | gibi | mriedem: OK, thats a different failure than the what originally was pasted to the bug report. I continue digging... | |
| 15:13:19 | cdent | another possibility is that uwsgi is (somehow, who knows) letting things get out of order | |
| 15:13:29 | mriedem | dansmith: do we log anything specific in that case? | |
| 15:13:40 | mriedem | i see a buttload of the "we're on a pike compute with all pike computes, so not healing allocations" all the time | |
| 15:13:41 | dansmith | mriedem: placement will log it | |
| 15:13:51 | dansmith | mriedem: okay then that probably means it's not | |
| 15:13:59 | mriedem | but ^ is from the periodic | |
| 15:14:18 | mriedem | there are paths in the RT that go into that code w/o consciously passing the has_ocata_computes flag | |
| 15:14:22 | mriedem | but i think it defaults to False anyway | |
| 15:15:13 | cdent | mriedem: have you got a set of logs you can make available? | |
| 15:15:18 | openstackgerrit | OpenStack Proposal Bot proposed openstack/os-vif stable/pike: Updated from global requirements https://review.openstack.org/493146 | |
| 15:16:02 | cdent | this doesn’t feel like something it’s going to be easy to reason about without some files to grep | |
| 15:16:48 | dansmith | cdent: it should be pretty easy to reproduce (or not) in a devstack and then you can instrument the code as needed | |
| 15:16:58 | openstackgerrit | OpenStack Proposal Bot proposed openstack/python-novaclient stable/pike: Updated from global requirements https://review.openstack.org/493187 | |
| 15:17:12 | mriedem | cdent: no, it's all local | |
| 15:17:19 | mriedem | well, in this devstack vm which is not local | |
| 15:17:28 | mriedem | but yeah i have the local.conf for the devstack if you want to reproduce | |
| 15:17:32 | cdent | mriedem: sure, but you have tar and such? | |
| 15:17:41 | mriedem | yeah | |
| 15:18:16 | mriedem | is there a standard way to tar up the journald logs? | |
| 15:18:33 | cdent | balls, I forgot about journald, meh | |
| 15:18:46 | mriedem | it's tar'ed up in devstack-gate | |
| 15:18:52 | mriedem | so i can just copy whatever we do in CI | |
| 15:19:46 | openstackgerrit | Rodolfo Alonso Hernandez proposed openstack/nova master: Change 'InstancePCIRequest' spec field https://review.openstack.org/449257 | |
| 15:19:47 | cdent | I can’t really look with any real attention until about 3 hours from now, but if you get a chance to do it, that’s great it will useful, if not, just the local.conf will do | |
| 15:19:57 | mriedem | https://github.com/openstack-infra/devstack-gate/blob/master/functions.sh#L698-L724 | |
| 15:21:25 | mriedem | you know, i could just do this with a devstack patch | |
| 15:21:27 | mriedem | that's easier | |
| 15:21:34 | mriedem | let the ci do the work | |
| 15:25:19 | mriedem | needless to say, i'm doing a terrible job of reviewing code or specs, or writing specs | |
| 15:26:01 | mriedem | sdague: can you get this stable/pike novaclient bug fix backport? https://review.openstack.org/#/c/495901/ | |
| 15:26:12 | mriedem | pretty nasty and we need to release it | |
| 15:26:18 | sdague | mriedem: looking | |
| 15:26:55 | sdague | +A | |
| 15:27:32 | dansmith | mriedem: speaking of reviewing, I'm not sure what else to do on the base switchover patch since the difference is not measurable on my box. I'm poking at the fault thing in the later patch, but we can just strip that out and work on it in parallel | |
| 15:27:52 | mriedem | i haven't reviewed the actual change yet | |
| 15:28:04 | mriedem | was just getting 1000 active vms locally to test the fault thing we talked about last night | |
| 15:28:14 | mriedem | i agree that what we found last night, for numbers, isn't worth holding things up | |
| 15:28:28 | mriedem | sdague: thanks | |
| 15:28:34 | dansmith | mriedem: okay, like I said in the etherpad, I don't think we're doing the fault thing in that patch yet | |
| 15:28:47 | dansmith | and applying the one that does it definitely has an impact | |
| 15:29:15 | mriedem | negative impact? | |
| 15:29:37 | dansmith | yeah | |
| 15:30:05 | mriedem | because we've added a new unconditional join i suppose | |
| 15:30:19 | dansmith | it's not a new join, but it's a new query yeah | |
| 15:30:25 | mriedem | yeah, was just thinking that | |
| 15:30:43 | dansmith | I'm messing with the later patch though to see if I can squash it out | |
| 15:30:45 | mriedem | like you said, we could plumb that in the db api | |
| 15:30:53 | openstackgerrit | OpenStack Proposal Bot proposed openstack/os-vif stable/newton: Updated from global requirements https://review.openstack.org/373293 | |
| 15:31:05 | dansmith | that's what I'm doing yeah | |
| 15:31:05 | mriedem | if faults is in expected_attrs, get the instances in deleted/error and include the faults on those? | |
| 15:32:04 | dansmith | yes | |
| 15:36:47 | cdent | do dansmith and mriedem have an exciting new etherpad they are willing to share? | |
| 15:37:06 | dansmith | cdent: https://etherpad.openstack.org/p/nova-instance-list | |
| 15:37:12 | dansmith | cdent: but that's just incidental to him finding the problem | |
| 15:37:21 | dansmith | it's not about the thing we were describing | |
| 15:37:35 | cdent | you know I’m a total junkie for all the info | |
| 15:37:55 | mriedem | started as getting a benchmark for testing before and after the big instance list change | |
| 15:38:10 | mriedem | but have also added todos for weird things we've seen, like the concurrent update detected thing with 500 instances | |
| 15:42:09 | mriedem | sdague: can i put [[post-config|$NOVA_CONF]] in stackrc? | |
| 15:42:44 | cdent | mriedem, dansmith, (and sean mooney): If any of you get a chance to look at the discussion on https://review.openstack.org/#/c/504540/ (limiting GET /allocation_candidates ) it’s gotten to the point where we are trying to decide what it is that we are actually optimizing for, so could do with more input | |
| 15:43:17 | sdague | mriedem: in local.conf | |
| 15:43:42 | mriedem | sdague: yeah, locally, but this is for running something through ci | |
| 15:44:51 | sdague | there isn't really anyway to jam it into stackrc, you need to do project-config changes for stuff like that | |
| 15:45:15 | sdague | if this is just for hacktastic stuff, just iniset whatever you want in lib/nova | |
| 15:45:16 | mriedem | ok, just checking, i can hack this other ways | |
| 15:45:21 | mriedem | yeah that's what i'm doing | |
| 15:55:01 | dansmith | mriedem: so with /servers/details on 2.53, we're pegging conductor real hard. With 2.1 there is no conductor interaction at all | |
| 15:55:13 | dansmith | so we must be doing something stupid in the later microversion | |
| 15:55:19 | dansmith | because we shouldn't be making rpc calls at all | |