| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-07-18 | |||
| 15:43:48 | jaypipes | zzzeek: thus me saying it was my mistake ;) | |
| 15:44:19 | gibi | jaypipes: I'm getting the same stacktrace after reodering the condition. git diff is on the top of the logs: http://paste.openstack.org/show/615749/ | |
| 15:44:20 | zzzeek | jaypipes: OK so, if you have "col <operator> othercol" that left/right is maintained | |
| 15:44:43 | zzzeek | jaypipes: also, it....shouldnt matter? unless you're trying to hit an index on oracle | |
| 15:45:22 | zzzeek | jaypipes: == operator is commutative... | |
| 15:45:50 | jaypipes | gibi: :( ok, back to the drawing board. I really don't know why a KeyError is being raised there. the root provider ID should be in the summaries dict since _get_usages_by_rp_and_rc() should be returning a record for that rp | |
| 15:46:12 | gibi | jaypipes: is there any log I can turn on to help? | |
| 15:46:34 | gibi | jaypipes: or if you provide a patch with extra LOGs then I can apply that | |
| 15:46:35 | jaypipes | zzzeek: right, but I was thinking maybe SA saw the == operator column order and maybe made the expression b LEFT JOIN a instead of a LEFT JOIN b. | |
| 15:46:41 | jaypipes | zzzeek: apparently not, though | |
| 15:46:55 | jaypipes | gibi: I'll do the latter | |
| 15:47:25 | gibi | jaypipes: OK. I'm still around for an hour or so then I can continue tomorrow | |
| 15:47:32 | gibi | jaypipes: thank again for helping | |
| 15:47:38 | jaypipes | gibi: I'll have a patch up in 5 mins. | |
| 15:47:44 | zzzeek | jaypipes: ah. no way :) | |
| 15:47:55 | gibi | jaypipes: I will test that! | |
| 15:48:26 | openstackgerrit | Matt Riedemann proposed openstack/nova master: DNM: test new style cinder attach with upgrades https://review.openstack.org/484860 | |
| 15:48:27 | mriedem | mgiles: thanks, testing it here ^ | |
| 15:49:41 | mgiles | miredem: great! thanks | |
| 15:49:56 | mgiles | mriedem ^ | |
| 15:51:03 | melwitt | mriedem: I think we've glossed over that in the review and the thought was, we can't count those atomically together with instances because they're in the API DB ... | |
| 15:51:26 | openstackgerrit | Jay Pipes proposed openstack/nova master: TESTING - DO NOT MERGE https://review.openstack.org/484862 | |
| 15:51:30 | jaypipes | gibi: ^^ | |
| 15:53:30 | jaypipes | dtantsur: looking at that devstack custom RCs patch now... | |
| 15:53:41 | mriedem | melwitt: do we count instances atomically across multiple cells? | |
| 15:53:42 | dtantsur | thanks! | |
| 15:53:52 | gibi | jaypipes: looking... | |
| 15:53:55 | mriedem | melwitt: don't we have a separate session for each cell db? | |
| 15:53:59 | melwitt | mriedem: no, each cell is atomic | |
| 15:54:17 | dansmith | mriedem: we can't count atomically across cells | |
| 15:55:03 | mriedem | right, | |
| 15:55:14 | mriedem | so my point is, saying we can't do it atomically b/c of the api db is kind of a cop out | |
| 15:55:25 | mriedem | since we can't do it atomically for multiple cells either | |
| 15:55:29 | dansmith | well, | |
| 15:55:34 | dansmith | but the cells don't overlap with each other | |
| 15:55:40 | mriedem | true, yes | |
| 15:55:41 | dansmith | the api db and the cell dbs do overlap | |
| 15:55:44 | mriedem | the race window is a problem | |
| 15:55:47 | mriedem | between build requests and instances | |
| 15:56:12 | mriedem | again, i don't think i'm advocating trying to account for build requests in flight at the same time as counting the instances | |
| 15:56:21 | mriedem | i just want to make sure we thought about it and are ok with it | |
| 15:57:54 | jaypipes | dtantsur: answered. | |
| 15:58:09 | melwitt | mriedem: yeah, I mean, I was thinking ideally we should count them because they are parts of instances | |
| 15:58:46 | dtantsur | jaypipes: thanks! I wonder if we should enroll nodes after nova-compute is started then | |
| 15:58:58 | dtantsur | jaypipes: to better emulate how things work in actual production | |
| 15:59:00 | dtantsur | wdyt? | |
| 15:59:03 | mriedem | the "for x in 5; nova boot --min-count 20...." worries me | |
| 15:59:39 | jaypipes | dtantsur: well, they automatically get enrolled when the nova-compute node starts, but it doesn't look like that has run by the time the test tries to boot an instance. | |
| 16:00:04 | dtantsur | this is suspicious.. okay, thanks for the hints. I'll take a deeper look tomorrow | |
| 16:00:17 | jaypipes | dtantsur: looks to be just a simple ordering issue to me. | |
| 16:00:37 | jaypipes | dtantsur: I will look further into it as well and leave comments on the patch. | |
| 16:00:45 | dtantsur | thanks :) | |
| 16:14:55 | melwitt | mriedem: yeah. the bad thing about counting both is it could get too much usage if you say, count build requests and then count instances, some of those instances could have been build requests a split second ago, and then you get too much usage after you add them together | |
| 16:17:16 | gibi | jaypipes: here is the log http://paste.openstack.org/show/615753/ | |
| 16:20:21 | jaypipes | gibi: excellent, that helps a lot, thank you! | |
| 16:20:40 | jaypipes | gibi: the usage information isn't being returned for rp 1 for some reason. | |
| 16:21:37 | gibi | jaypipes: rp 1 is the compute provider? | |
| 16:21:48 | jaypipes | gibi: yeah | |
| 16:22:11 | melwitt | mriedem, dansmith: apparently sqlalchemy supports two-phase commit. do you think that's something we could use to get atomic across databases? http://docs.sqlalchemy.org/en/latest/orm/session_transaction.html#enabling-two-phase-commit | |
| 16:22:41 | jaypipes | melwitt: I wouldn't go that route. It's a pain in the ass, frankly. | |
| 16:23:05 | dansmith | not worth it, IMHO | |
| 16:24:00 | melwitt | hm, okay | |
| 16:24:10 | jaypipes | gibi: what's the result of SELECT * FROM inventories WHERE resource_provider_id = 1? | |
| 16:26:08 | melwitt | dansmith: I guess the rechecking quota stuff takes care of the gap between build requests and instances ... in that after creating the instance objects we check quota again. so maybe that's how not counting build requests can be fine (if someone is configured to be strict about quotas) | |
| 16:26:26 | dansmith | melwitt: yep | |
| 16:26:32 | melwitt | mriedem: ^ | |
| 16:26:56 | melwitt | whew, good. | |
| 16:27:33 | openstackgerrit | Merged openstack/nova master: [placement] fix 500 error when allocating to bad class https://review.openstack.org/484162 | |
| 16:27:40 | mriedem | melwitt: you mean the recheck performed in conductor? | |
| 16:27:48 | melwitt | mriedem: yeah | |
| 16:28:05 | gibi | jaypipes: http://paste.openstack.org/show/615755/ | |
| 16:28:33 | mriedem | yeah, my point earlier today was the difference in ux - before counting qoutas, you'd go overquota and fail the api with a 403 and no instances are created - after we if we catch it in conductor, the instances all get put into ERROR state and you have to clean them up | |
| 16:28:38 | mriedem | melwitt: not the end of the world, but it's different | |
| 16:29:32 | dansmith | now I'm forgetting, | |
| 16:29:45 | dansmith | but the first check is done in api, and the extra if configured check is in conductor, right? | |
| 16:29:57 | melwitt | dansmith: yes | |
| 16:29:58 | mriedem | yeah | |
| 16:30:03 | mriedem | we recheck for cells v1 in the api | |
| 16:30:06 | dansmith | right, so the ux change is only across the boundary | |
| 16:30:07 | mriedem | b/c that's where the instances are created | |
| 16:30:10 | mriedem | but...cellsv1 | |
| 16:30:11 | dansmith | which isn't that big of a deal, IMHO | |
| 16:30:41 | dansmith | quota is so leaky right now, | |
| 16:30:41 | mriedem | it's not the end of the world, | |
| 16:30:57 | mriedem | i only ever worry about the NFV robots hammering things in unexpected ways | |
| 16:30:58 | jaypipes | gibi: are you calling GET /allocation_candidates?resources=CUSTOM_MAGIC:512 or are you including VCPU and MEMORY_MB in the request as well? | |
| 16:31:08 | dansmith | we could have added another check to the current stuff where we decided to fail your instance because we healed the quota and the numbers no longer work out | |
| 16:31:32 | openstackgerrit | Merged openstack/nova master: Support tag instances when boot(4/4) https://review.openstack.org/469800 | |
| 16:31:38 | melwitt | every quota.reserve() call heals before checking | |
| 16:31:54 | mriedem | heal == refresh()? | |
| 16:31:55 | melwitt | er, it heals YOUR quota. but not others in your project. so you could still be screwed | |
| 16:32:26 | dansmith | melwitt: it doesn't heal negative things, right? | |
| 16:32:42 | dansmith | like when things get all out of whack and require admin and nova-manage intervention | |
| 16:32:45 | dansmith | that's what I meant | |
| 16:32:49 | gibi | jaypipes: I'm calling just MAGIC:512 first, that results in HTTP 500 then I call with both MAGIC and VCPU + MEMORY that goes through but results in an empty response | |
| 16:32:51 | melwitt | I don't think it stores negative usage | |
| 16:33:23 | jaypipes | gibi: can you paste the logs from the second request for VCPU, MAGIC and MEMORY_MB? | |
| 16:33:24 | melwitt | yeah, you need nova-manage when not all users in the cloud are active in booting instances and their quota is out of whack and affecting those who are active | |
| 16:33:41 | dansmith | melwitt: that's all I mean.. when the current broke-ass shit gets broke-ass | |
| 16:33:47 | melwitt | yeah | |