| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-29 | |||
| 17:57:09 | mriedem | i think, whenever we did the double up thing in the scheduler anyway | |
| 17:58:08 | mriedem | oh sorry the doubled up allocations happened the day of FF | |
| 17:58:31 | mriedem | that can't be right | |
| 17:58:52 | mriedem | oh yeah no it was FF | |
| 17:58:53 | mriedem | :( | |
| 17:59:15 | dansmith | mriedem: move ops in general yeah | |
| 18:00:02 | dansmith | I'll be pushing some more stuff up in that series in a bit | |
| 18:00:16 | dansmith | gonna try to break things into really small bits where possible per my usual, | |
| 18:00:34 | dansmith | but hopefully to make each change a clear and understandable win | |
| 18:00:42 | mriedem | ok, | |
| 18:00:51 | mriedem | i'm checking out the unshelve failure flows to see what we might have missed | |
| 18:01:00 | mriedem | and then will start working on the evacuate bug later | |
| 18:02:35 | mriedem | i guess i should start a retrospective etherpad for pike before the ptg... | |
| 18:02:44 | mriedem | i'm not sure i want to even think about what we did wrong | |
| 18:02:57 | cdent | some stuff. we admit it. done | |
| 18:02:58 | mriedem | that's reserved for when i wake up at 2am | |
| 18:11:56 | mriedem | ok https://etherpad.openstack.org/p/nova-pike-retrospective | |
| 18:11:59 | mriedem | posting to ML | |
| 18:13:10 | tomtomtom | hello, I'm having trouble launching instances with ephemeral volumes via ceph. anyone got any docs or pointers for such a configuration? | |
| 18:13:47 | tomtomtom | cinder, nova, and ceph "seem" to be creating the volume but nova comes up with "no bootable device" each time. | |
| 18:15:47 | mriedem | tomtomtom: https://docs.openstack.org/nova/latest/user/block-device-mapping.html ? | |
| 18:18:46 | sean-k-mooney | tomtomtom: does normal booting work wtih ceph backed root device or only fails with ephemeral disk | |
| 18:19:22 | mriedem | tomtomtom: can you clarify what you mean by 'ephemeral volumes'? | |
| 18:19:27 | mriedem | are you actually booting from volume? | |
| 18:19:45 | mriedem | and you consider it ephemeral because delete_on_termination=True? | |
| 18:21:06 | sean-k-mooney | alternitvly do you mean you have allocated ephmeral storage in the flavor and have confiuged nova to back all vm storage with ceph volumes | |
| 18:27:28 | mriedem | looks like we're ok wrt allocations during unshelve | |
| 18:27:39 | mriedem | conductor calls scheduler to pick a host, creates the allocations, and casts to compute | |
| 18:27:44 | mriedem | where the instance claim happens | |
| 18:27:57 | mriedem | no retries | |
| 18:28:15 | mriedem | MAYBE WE SHOULD BUILD RETRY LOGIC INTO UNSHELVE | |
| 18:29:15 | mriedem | actually, :) | |
| 18:29:26 | sean-k-mooney | cdent: lol | |
| 18:29:59 | sean-k-mooney | cdent: you realise that that would make it your problem to fix | |
| 18:30:52 | mriedem | we make the claim, and then try to spawn the instance, if that fails we unset the instance.host/node values, | |
| 18:30:57 | mriedem | but i'm not sure that we remove the allocations | |
| 18:31:50 | cdent | b@ll$ | |
| 18:32:18 | mriedem | yeah we don't | |
| 18:33:01 | mriedem | we'd abort the claim on the exit of this https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/manager.py#L4485 | |
| 18:33:22 | mriedem | https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/resource_tracker.py#L414 | |
| 18:33:33 | cdent | mriedem: maybe you need to reset your scanning algorithm: look for where we do, because we started from the point of not thinking about it, thus... | |
| 18:33:46 | mriedem | that method, by default, sets has_ocata_computes=False | |
| 18:33:46 | mriedem | https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/resource_tracker.py#L1012 | |
| 18:33:59 | mriedem | which means we won't fix the allocations https://github.com/openstack/nova/blob/16.0.0.0rc2/nova/compute/resource_tracker.py#L1045 | |
| 18:34:17 | mriedem | cdent: we just always relied on the periodic to heal things | |
| 18:34:29 | mriedem | which is the same code that would have done this before ^ | |
| 18:34:48 | cdent | how about we just put it back, for now? | |
| 18:35:07 | mriedem | because we want to remove it so the RT isn't trampling over thigns | |
| 18:35:23 | mriedem | it just really means that we have to be very explicity about dealing with allocations everywhere | |
| 18:35:44 | mriedem | but, that's probably for the best in the long run | |
| 18:36:12 | cdent | then in that case there’s no reason to express surprise that we aren’t handling things, yeah? | |
| 18:36:27 | mriedem | correct | |
| 18:40:00 | cdent | the functional tests are great, but fairly heavy | |
| 18:40:16 | cdent | s/heavy/cumbersome to create | |
| 18:43:37 | mriedem | https://bugs.launchpad.net/nova/+bug/1713796 | |
| 18:43:38 | openstack | Launchpad bug 1713796 in OpenStack Compute (nova) "Failed unshelve does not remove allocations from destination node" [High,Triaged] | |
| 18:43:51 | mriedem | there are at least 2 ways unshelve can fail there which we don't cleanup the allocations | |
| 18:45:03 | mriedem | maybe we need an undo_allocations decorator for several methods in the compute manager | |
| 18:46:41 | cdent | are the exit conditions workable for a decorator (or, to put it another way, how is that different from the two other ideas above?) | |
| 18:47:32 | mriedem | it could be messy in a decorator, probably lots of conditional logic based on the operation being performed | |
| 18:47:37 | mriedem | which is based on the task_state | |
| 18:48:03 | mriedem | so i'm not going to bother thinking about that for now | |
| 18:49:06 | mriedem | regarding the functional tests, i think we need those regardless | |
| 18:49:16 | cdent | oh, yeah, I wasn’t saying we should get rid of them | |
| 18:49:23 | mriedem | since we didn't have much in the way of negative functional scenario tests that were asserting allocations | |
| 18:49:42 | cdent | but I was wondering if something like a shell script that did curl based validations could be used to confirm where the bugs are | |
| 18:50:15 | cdent | nova < some command> […] curl to verify some allocation | |
| 18:54:13 | sean-k-mooney | cdent: why curl vs adding support for placement to the openstack client | |
| 18:54:21 | sean-k-mooney | we will need that evenutally anyway | |
| 18:54:35 | sean-k-mooney | we will also need it for traits | |
| 18:54:47 | mriedem | sean-k-mooney: there are patches for an osc placement plugin | |
| 18:55:05 | mriedem | https://review.openstack.org/#/q/project:openstack/osc-placement | |
| 18:55:18 | mriedem | cdent: yeah i thought about something like that the other night, | |
| 18:55:30 | mriedem | like, we have a post test hook into our CI environments to do things | |
| 18:55:31 | sean-k-mooney | oh cool then ya that would be easier to use in test code then curl as it handels gettin ghte keysone auth tokens for you which is a pain | |
| 18:55:46 | mriedem | cdent: assuming we cleaned everything up properly, the nodes shouldn't have any allocations against them at the end of the job | |
| 18:56:44 | mriedem | i think we run the archive instances stuff through that in our jobs somewhere | |
| 18:57:43 | sean-k-mooney | mriedem: i think we do someting similar in our thridpart collecd ci | |
| 18:57:44 | mriedem | dansmith: don't we run archive_deleted_rows in our ci somewhere? | |
| 18:58:32 | mriedem | yeah the post_test_hook.sh, i bet project-config is busted again | |
| 18:58:45 | mriedem | nova/tools/hooks/post_test_hook.sh | |
| 18:59:23 | sean-k-mooney | soory got confued ignore the collectd comment but what i was goning to say is im not sure if post_test_hook.sh runs after the logs are upladed or not | |
| 18:59:43 | mriedem | oh it's only in the nova-next job | |
| 19:00:05 | mriedem | kablam! http://logs.openstack.org/61/496861/1/check/gate-tempest-dsvm-neutron-nova-next-full-ubuntu-xenial-nv/572ceed/logs/devstack-gate-post_test_hook.txt.gz | |
| 19:00:13 | cdent | sean-k-mooney: I was thinking curl because of something quick and dirty that wouldn’t necessarily be a permanently running test, rather a diagnostic tool for nowish | |
| 19:00:32 | mriedem | cdent: so it would be cool to write something that queries for allocations in the post test hook at the end of a job to see if they are 0 or not | |
| 19:00:33 | cdent | also, what’s hard about getting a token? | |
| 19:00:59 | mriedem | anyway, i'll post the idea to the ML and people can think about it | |
| 19:01:04 | sean-k-mooney | cdent: more you just need to do that first befor just calling curl and add it to the correct header | |
| 19:01:25 | cdent | mriedem: would we’d need to track instance ids throughout the test or can we simply get all instances, including deleted ones? | |
| 19:01:32 | sean-k-mooney | cdent: if you want something quick just run a mysql command directly against the db | |
| 19:02:04 | sean-k-mooney | if there are still allocations in the output you know something is messed up | |
| 19:02:09 | cdent | sean-k-mooney: I tend do to a lot of export TOKEN=$(openstack token issue …); curl blah | |
| 19:02:40 | sean-k-mooney | cdent: huh i didnt know you could do openstack token issue ... | |
| 19:02:42 | cdent | but yeah, a db query would be fine too, but…boring :) | |
| 19:02:49 | sean-k-mooney | that makes things alost eaisier | |
| 19:03:10 | cdent | also it would mean that we are assuming that placement will always have a db backend ;) | |
| 19:03:19 | sean-k-mooney | any time i have used curl like that i have manually used curl to get the token by calling keysone myself which is a pain | |
| 19:03:28 | cdent | ouch. yeah. that is painful | |
| 19:05:32 | sean-k-mooney | cdent: and yes i am assiming a db backend. honestly i spend enough time getting our ci test to stop going to the db directly and use the api instead so ya curl would be better if its going to last long enough for db assumtions to expire | |