| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-06-08 | |||
| 18:32:22 | mgagne | I'm sure yours are more numerous than mines | |
| 18:33:00 | mgagne | alright, will go back to my ffu dungeon then :D | |
| 18:34:27 | mriedem | say hi to lyaaaaaaaarwood for me | |
| 18:35:40 | mgagne | is it a friend or some beast I should find and fight in the dungeon? :D | |
| 18:36:18 | mgagne | oh, a ffu expert :D | |
| 18:36:47 | openstackgerrit | Chris Dent proposed openstack/nova stable/queens: Ensure resource class cache when listing usages https://review.openstack.org/573811 | |
| 18:38:18 | mriedem | lyaaaaaaaarwood is a level 12 FFU paladin with a ring of +3 "soft speaking" | |
| 18:46:17 | mriedem | cfriesen: https://review.openstack.org/573813 | |
| 18:49:18 | mriedem | cfriesen: openstack server migration list / abort etc is probably good yeah | |
| 18:49:21 | mriedem | since we have openstack server migrate | |
| 18:49:29 | mriedem | and don't want to get confused with openstack volume migrate | |
| 18:49:37 | mriedem | i leave stuff like that up to the osc ux wizards | |
| 20:12:19 | jroll | is it expected that we should be able to run 'show' or 'delete' via openstackclient on instances in cell0? | |
| 20:12:22 | jroll | using ocata | |
| 20:12:50 | jroll | even list bombs out | |
| 20:13:19 | jroll | http://paste.openstack.org/show/6YrSmjMSo0lIxyFjbPIz/ | |
| 20:13:35 | jroll | looks like it's hitting the wrong database when trying to refresh the instance in _load_flavor() | |
| 20:14:21 | jroll | looks like this bug which expired: https://bugs.launchpad.net/nova/+bug/1749167 | |
| 20:14:23 | openstack | Launchpad bug 1749167 in OpenStack Compute (nova) "nova show can not get an instance information, and this instance can be queried from nova list." [Undecided,Expired] | |
| 20:14:42 | jroll | by the way bowser was trying to ask about the actual scheduling problem, sounds like maybe this is expected? | |
| 20:17:21 | melwitt | jroll: I don't think it's expected | |
| 20:18:22 | melwitt | we had bugs around instance list/show back then, which we fixed. I'm looking through bugs to see if any were this | |
| 20:18:27 | temka | The fact that it's hitting it in flavor... | |
| 20:18:41 | jroll | I didn't immediately see it searching everything for "cell0" | |
| 20:18:43 | melwitt | you have the latest ocata or an earlier one? | |
| 20:18:52 | jroll | should be latest, lemme verify | |
| 20:18:54 | melwitt | yeah, I'm not immediately finding anything either | |
| 20:19:03 | melwitt | it's hard to find these things | |
| 20:19:59 | jroll | well wtf, our latest upstream commit is 125dd1f30fdaf50182256c56808a5199856383c7 | |
| 20:20:04 | jroll | which was february | |
| 20:20:22 | melwitt | I'm not sure if it matters, just wanted to make sure I understand which code you have | |
| 20:20:49 | jroll | it matters that I have a broken assumption :) | |
| 20:22:24 | jroll | similar: https://github.com/openstack/nova/commit/e0c1d461af0701adb94e6974f363e12395ed0162 | |
| 20:22:42 | jroll | but I don't think it's the same | |
| 20:23:11 | temka | jroll, 2e68b2298e94a15d1282c0fb46804b9efa6c8b3a ? | |
| 20:23:30 | temka | Seems old tho | |
| 20:23:40 | jroll | yeah, we would have that | |
| 20:24:19 | melwitt | it's trying to lazy-load the flavor from the database, first thing it does while doing that is the lookup the instance, and that is not found because it's looking in the wrong database | |
| 20:24:29 | jroll | exactly | |
| 20:24:32 | melwitt | this sounds so familiar, just not finding a bug that matches yet | |
| 20:24:41 | temka | melwitt, yep, same here | |
| 20:24:46 | temka | I swear I've seen this | |
| 20:25:40 | jroll | melwitt: in the meantime, we can manually delete it from cell0.instances and nova_api.instance_mappings, trigger a quota reset... will all the other tables get cleaned up on delete? | |
| 20:28:10 | melwitt | I will say that ocata is particularly fraught with problems because the way we were doing cell targeting back then had some fundamental problems which we changed in pike, but was a huge change that backporting would involve backporting a ton of disjointed things in pike, so we abandoned it | |
| 20:28:57 | jroll | mmm | |
| 20:28:59 | melwitt | (assuming that things are okay-enough in ocata. if they're not, then we have to find another way to solve it) | |
| 20:29:53 | melwitt | jroll: yeah, I think that would do it as a manual cleanup. are you seeing this on/during an upgrade or on already-upgraded-and-been-running clusters? | |
| 20:30:24 | jroll | melwitt: already upgraded | |
| 20:30:39 | jroll | we shut down the API for this upgrade | |
| 20:30:42 | melwitt | jroll: okay, so any new instance that fails to schedule gets in this un-listable state | |
| 20:31:02 | jroll | trying to verify that now | |
| 20:31:10 | jroll | this instance is still in BUILD, rather than ERROR | |
| 20:31:19 | jroll | and is the only thing in instance_mappings with cell_id=NULL | |
| 20:31:44 | melwitt | oh, that's interesting. but yet it's in cell0, in BUILD state but not ERROR | |
| 20:32:17 | jroll | yep | |
| 20:32:22 | jroll | and built > 16 hours ago | |
| 20:32:28 | melwitt | cell_id=NULL is a problem. that should be cell0 if it's in cell0 | |
| 20:32:32 | melwitt | hmm | |
| 20:32:40 | jroll | task_state=scheduling | |
| 20:32:51 | jroll | so it's like it got dropped by rabbit or something? idk | |
| 20:33:53 | melwitt | yeah, I dunno. we have also heard from CERN folks of some instance_mappings ending up with cell_id=NULL but I've not understood the conditions under which it happened | |
| 20:35:21 | melwitt | while instances are in the middle of scheduling, they are supposed to be listed/shown using their build_requests in the nova_api database, IIRC | |
| 20:36:07 | jroll | hm, can't run server show on other instances in cell0 either | |
| 20:36:12 | jroll | so maybe that's unrelated | |
| 20:36:21 | jroll | though that 404s instead of 500s | |
| 20:37:58 | jroll | cell_id=1 doesn't break server list either, but doesn't show up in the response | |
| 20:38:24 | melwitt | okay, first check the db connection string for cell0 in nova_api.cell_mappings and make sure it's right, then look in nova_api.instance_mappings and check if cell_id is cell0, then check if the instance is in nova_cell0.instances. that's how it's all connected | |
| 20:39:35 | jroll | yeah, that's all correct | |
| 20:39:44 | melwitt | it will go instance_mappings.cell_id => db connect to cell_mappings.cell_id database_connection string => <cell database>.instances (if the instance has completed scheduling, either via success or error) | |
| 20:40:17 | melwitt | if it did not complete scheduling, it will live in build_requests only and have cell_id=NULL, I believe | |
| 20:40:42 | jroll | nod | |
| 20:40:50 | melwitt | :/ okay, if that's all correct I don't understand how it could _not_ be able to 'nova show' it | |
| 20:41:04 | melwitt | does that still fail with the flavor lazy-load thing? | |
| 20:41:13 | jroll | digging logs now for the 404 | |
| 20:41:45 | melwitt | meaning, the instance that checks out as far as all the things connected properly, fails on the flavor lazy-load? | |
| 20:44:20 | jroll | right, so the one that is properly mapped to cell0 | |
| 20:44:53 | jroll | is throwing a 404 | |
| 20:45:20 | jroll | HTTP exception thrown: Instance 75e193c5-6700-4d67-be44-7020fb4129ed could not be found. | |
| 20:45:34 | melwitt | okay | |
| 20:46:28 | melwitt | still trying to parse whether that lazy-load read_deleted bug you linked earlier is this or not | |
| 20:46:48 | melwitt | the bug seems light on details | |
| 20:46:58 | jroll | yeah, ditto | |
| 20:48:48 | melwitt | okay, so this is about deleted instances not being found when trying to lazy-load other instance properties. that seems kinda chicken and egg. (how did you get the deleted instance to lazy-load from in the first place?) | |
| 20:49:15 | melwitt | but your instances have not been deleted, so I think it's probably not the same bug | |
| 20:51:37 | melwitt | jroll: what's your service version for the nova-osapi_compute service? | |
| 20:52:09 | melwitt | probably in nova_cell0.services | |
| 20:52:49 | jroll | right, not deleted | |
| 20:54:10 | jroll | debug logs gave me jack | |
| 20:54:17 | jroll | no records in nova_cell0.services O_o | |
| 20:54:34 | melwitt | must be in a different db then, I don't know where to expect it | |
| 20:55:13 | melwitt | oh, maybe just nova? | |
| 20:55:20 | jroll | 16 in nova yeah | |
| 20:55:25 | melwitt | the non-compute service records are all together. okay | |
| 20:55:56 | jroll | bah, dogs need to go out, bbiab | |
| 20:56:02 | melwitt | k | |
| 20:56:39 | melwitt | so you should be in here https://github.com/openstack/nova/blob/stable/ocata/nova/compute/api.py#L2322-L2347 | |
| 21:00:25 | melwitt | oh, wait, it never even gets there. uuugggh | |
| 21:00:56 | melwitt | this is trying to look up the instance long before instance mappings are even begun to be examined | |
| 21:03:55 | jroll | the "bad" one does get that far | |
| 21:04:31 | jroll | but can't tell on the "good" one since there's no traceback or anything on the 404 | |