| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-29 | |||
| 18:19:27 | mriedem | are you actually booting from volume? | |
| 18:19:45 | mriedem | and you consider it ephemeral because delete_on_termination=True? | |
| 18:21:06 | sean-k-mooney | alternitvly do you mean you have allocated ephmeral storage in the flavor and have confiuged nova to back all vm storage with ceph volumes | |
| 18:27:28 | mriedem | looks like we're ok wrt allocations during unshelve | |
| 18:27:39 | mriedem | conductor calls scheduler to pick a host, creates the allocations, and casts to compute | |
| 18:27:44 | mriedem | where the instance claim happens | |
| 18:27:57 | mriedem | no retries | |
| 18:28:15 | mriedem | MAYBE WE SHOULD BUILD RETRY LOGIC INTO UNSHELVE | |
| 18:29:15 | mriedem | actually, :) | |
| 18:29:26 | sean-k-mooney | cdent: lol | |
| 18:29:59 | sean-k-mooney | cdent: you realise that that would make it your problem to fix | |
| 18:30:52 | mriedem | we make the claim, and then try to spawn the instance, if that fails we unset the instance.host/node values, | |
| 18:30:57 | mriedem | but i'm not sure that we remove the allocations | |
| 18:31:50 | cdent | b@ll$ | |
| 18:32:18 | mriedem | yeah we don't | |
| 18:33:01 | mriedem | we'd abort the claim on the exit of this https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/manager.py#L4485 | |
| 18:33:22 | mriedem | https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/resource_tracker.py#L414 | |
| 18:33:33 | cdent | mriedem: maybe you need to reset your scanning algorithm: look for where we do, because we started from the point of not thinking about it, thus... | |
| 18:33:46 | mriedem | https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/resource_tracker.py#L1012 | |
| 18:33:46 | mriedem | that method, by default, sets has_ocata_computes=False | |
| 18:33:59 | mriedem | which means we won't fix the allocations https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/resource_tracker.py#L1045 | |
| 18:34:17 | mriedem | cdent: we just always relied on the periodic to heal things | |
| 18:34:29 | mriedem | which is the same code that would have done this before ^ | |
| 18:34:48 | cdent | how about we just put it back, for now? | |
| 18:35:07 | mriedem | because we want to remove it so the RT isn't trampling over thigns | |
| 18:35:23 | mriedem | it just really means that we have to be very explicity about dealing with allocations everywhere | |
| 18:35:44 | mriedem | but, that's probably for the best in the long run | |
| 18:36:12 | cdent | then in that case there’s no reason to express surprise that we aren’t handling things, yeah? | |
| 18:36:27 | mriedem | correct | |
| 18:40:00 | cdent | the functional tests are great, but fairly heavy | |
| 18:40:16 | cdent | s/heavy/cumbersome to create | |
| 18:43:37 | mriedem | https://bugs.launchpad.net/nova/+bug/1713796 | |
| 18:43:38 | openstack | Launchpad bug 1713796 in OpenStack Compute (nova) "Failed unshelve does not remove allocations from destination node" [High,Triaged] | |
| 18:43:51 | mriedem | there are at least 2 ways unshelve can fail there which we don't cleanup the allocations | |
| 18:45:03 | mriedem | maybe we need an undo_allocations decorator for several methods in the compute manager | |
| 18:46:41 | cdent | are the exit conditions workable for a decorator (or, to put it another way, how is that different from the two other ideas above?) | |
| 18:47:32 | mriedem | it could be messy in a decorator, probably lots of conditional logic based on the operation being performed | |
| 18:47:37 | mriedem | which is based on the task_state | |
| 18:48:03 | mriedem | so i'm not going to bother thinking about that for now | |
| 18:49:06 | mriedem | regarding the functional tests, i think we need those regardless | |
| 18:49:16 | cdent | oh, yeah, I wasn’t saying we should get rid of them | |
| 18:49:23 | mriedem | since we didn't have much in the way of negative functional scenario tests that were asserting allocations | |
| 18:49:42 | cdent | but I was wondering if something like a shell script that did curl based validations could be used to confirm where the bugs are | |
| 18:50:15 | cdent | nova < some command> […] curl to verify some allocation | |
| 18:54:13 | sean-k-mooney | cdent: why curl vs adding support for placement to the openstack client | |
| 18:54:21 | sean-k-mooney | we will need that evenutally anyway | |
| 18:54:35 | sean-k-mooney | we will also need it for traits | |
| 18:54:47 | mriedem | sean-k-mooney: there are patches for an osc placement plugin | |
| 18:55:05 | mriedem | https://review.openstack.org/#/q/project:openstack/osc-placement | |
| 18:55:18 | mriedem | cdent: yeah i thought about something like that the other night, | |
| 18:55:30 | mriedem | like, we have a post test hook into our CI environments to do things | |
| 18:55:31 | sean-k-mooney | oh cool then ya that would be easier to use in test code then curl as it handels gettin ghte keysone auth tokens for you which is a pain | |
| 18:55:46 | mriedem | cdent: assuming we cleaned everything up properly, the nodes shouldn't have any allocations against them at the end of the job | |
| 18:56:44 | mriedem | i think we run the archive instances stuff through that in our jobs somewhere | |
| 18:57:43 | sean-k-mooney | mriedem: i think we do someting similar in our thridpart collecd ci | |
| 18:57:44 | mriedem | dansmith: don't we run archive_deleted_rows in our ci somewhere? | |
| 18:58:32 | mriedem | yeah the post_test_hook.sh, i bet project-config is busted again | |
| 18:58:45 | mriedem | nova/tools/hooks/post_test_hook.sh | |
| 18:59:23 | sean-k-mooney | soory got confued ignore the collectd comment but what i was goning to say is im not sure if post_test_hook.sh runs after the logs are upladed or not | |
| 18:59:43 | mriedem | oh it's only in the nova-next job | |
| 19:00:05 | mriedem | kablam! http://logs.openstack.org/61/496861/1/check/gate-tempest-dsvm-neutron-nova-next-full-ubuntu-xenial-nv/572ceed/logs/devstack-gate-post_test_hook.txt.gz | |
| 19:00:13 | cdent | sean-k-mooney: I was thinking curl because of something quick and dirty that wouldn’t necessarily be a permanently running test, rather a diagnostic tool for nowish | |
| 19:00:32 | mriedem | cdent: so it would be cool to write something that queries for allocations in the post test hook at the end of a job to see if they are 0 or not | |
| 19:00:33 | cdent | also, what’s hard about getting a token? | |
| 19:00:59 | mriedem | anyway, i'll post the idea to the ML and people can think about it | |
| 19:01:04 | sean-k-mooney | cdent: more you just need to do that first befor just calling curl and add it to the correct header | |
| 19:01:25 | cdent | mriedem: would we’d need to track instance ids throughout the test or can we simply get all instances, including deleted ones? | |
| 19:01:32 | sean-k-mooney | cdent: if you want something quick just run a mysql command directly against the db | |
| 19:02:04 | sean-k-mooney | if there are still allocations in the output you know something is messed up | |
| 19:02:09 | cdent | sean-k-mooney: I tend do to a lot of export TOKEN=$(openstack token issue …); curl blah | |
| 19:02:40 | sean-k-mooney | cdent: huh i didnt know you could do openstack token issue ... | |
| 19:02:42 | cdent | but yeah, a db query would be fine too, but…boring :) | |
| 19:02:49 | sean-k-mooney | that makes things alost eaisier | |
| 19:03:10 | cdent | also it would mean that we are assuming that placement will always have a db backend ;) | |
| 19:03:19 | sean-k-mooney | any time i have used curl like that i have manually used curl to get the token by calling keysone myself which is a pain | |
| 19:03:28 | cdent | ouch. yeah. that is painful | |
| 19:05:32 | sean-k-mooney | cdent: and yes i am assiming a db backend. honestly i spend enough time getting our ci test to stop going to the db directly and use the api instead so ya curl would be better if its going to last long enough for db assumtions to expire | |
| 19:06:13 | sean-k-mooney | anyway time for food, see ye tomorow | |
| 19:07:27 | mriedem | cdent: i think we'd have to account for instances that aren't yet deleted | |
| 19:07:44 | mriedem | cdent: but, the post_test_hook stuff doesn't fail the job i don't think, so this would just be an early report to start | |
| 19:07:53 | mriedem | so we can kick ideas around | |
| 19:20:58 | mriedem | sdague: do you just want the commit message updated for this? https://review.openstack.org/#/c/457636/6/lib/placement | |
| 19:21:49 | sdague | mriedem: ok, this is long enough ago, I'm going to have to regain context | |
| 19:22:09 | sdague | I think the question is... do you always want that library installed? | |
| 19:22:34 | sdague | because, that won't install it in the normal case | |
| 19:22:41 | sdague | only when you specify it as LIBS_FROM_GIT | |
| 19:22:43 | mriedem | yes, we are going to want it in all cases i think | |
| 19:22:54 | sdague | which won't give you the release version, only the master version | |
| 19:22:57 | mriedem | right, | |
| 19:23:09 | sdague | ok, you you'll need an else | |
| 19:23:11 | mriedem | i think roman did that because he was working on getting an osc placement functional test job to run from a proposed change | |
| 19:23:21 | sdague | that does a pip_install osc-placement | |
| 19:23:41 | mriedem | ok. shouldn't the LIBS_FROM_GIT stuff be handled generically? | |
| 19:23:50 | mriedem | like, the job in project-config defines that i thought | |
| 19:24:02 | mriedem | and you had me do something like this for os-traits | |
| 19:24:07 | mriedem | to follow how the oslo libs are done | |
| 19:24:40 | mriedem | yeah, devstack/lib/libraries | |
| 19:25:33 | mriedem | https://github.com/openstack-dev/devstack/commit/aefc926cd45b2dc74d98f89e3a3b4cc92f2090ff | |
| 19:25:41 | mriedem | so i guess i can push something to do that for osc-placement | |
| 19:33:34 | mriedem | sdague: i should use pip_install_gr right? | |