| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-24 | |||
| 18:36:36 | mriedem | if you don't have ocata computes, that won't correct this now | |
| 18:36:38 | dansmith | which won't happen in pike land | |
| 18:36:49 | mriedem | we're not handling the case that prep_resize fails | |
| 18:36:57 | mriedem | and cleaning up the allocation | |
| 18:36:59 | mriedem | that the scheduler created | |
| 18:37:01 | dansmith | yeah, so you're talking about the case where we've doubled things in the scheduler and don't undouble them if we fail in prep right? | |
| 18:37:07 | mriedem | yup | |
| 18:37:38 | mriedem | so maybe this falls under the same bug i reported for when live migration fails and we don't cleanup | |
| 18:37:56 | dansmith | I would like to point out that if we were doing the allocation thing in the conductor instead of the scheduler, we'd have this all in an auto-cleanup context manager that would roll back the doubling if we failed to kick off a thing | |
| 18:37:58 | mriedem | this is essentially the same kind of fix probably, a periodic checking for failed migrations and cleaning up allocations related to them | |
| 18:38:20 | dansmith | well, we should clean up allocations any time we have a solid failure and know where the instance remains, | |
| 18:38:30 | dansmith | and a failure in prep is that case, right? we know we didn't move anything | |
| 18:39:06 | mriedem | prep_resize is a cast from conductor so i'm not sure how that would auto-cleanup in this case | |
| 18:39:24 | dansmith | it's a cast from conductor to compute? | |
| 18:39:34 | mriedem | yeah | |
| 18:39:41 | dansmith | ah, yeah, I see | |
| 18:39:57 | mriedem | so, remove the dest node allocation when not resizing to same host is simple | |
| 18:40:00 | dansmith | that's legacy from when the api was doing it I think, but.. | |
| 18:40:22 | mriedem | the resize to same host cleanup is shittier, since we basically need to overwrite the allocation back to the original flavor | |
| 18:40:37 | mriedem | which is basically just doing our RT overwrite stuff again, but in a different place | |
| 18:40:38 | dansmith | well, | |
| 18:40:51 | dansmith | we can just subtract what the new flavor would have had in it right? | |
| 18:41:07 | dansmith | not just regenerate, but subtract the new_flavor from our allocation if it's same host | |
| 18:41:17 | dansmith | merge_resources() with new_flavor and -1 as the sign | |
| 18:41:25 | dansmith | that will avoid trampling on shared things | |
| 18:41:30 | mriedem | sure | |
| 18:42:00 | mriedem | god i should probably have a test for the resize to same host case then too... | |
| 18:42:08 | mriedem | and it's nearly 2pm here | |
| 18:42:57 | mriedem | anyway, will work on the easy one first | |
| 18:44:47 | dansmith | mriedem: so you mean the fix for the thing you're testing in 497541 yes? | |
| 18:45:06 | mriedem | yeah | |
| 18:45:24 | mriedem | i don't know if we should do the resize to same host fix in here too or leave that for another bug | |
| 18:45:35 | mriedem | since it'd be a different test | |
| 18:46:00 | dansmith | if it doesn't really overlap then separate is probably best | |
| 18:46:03 | mriedem | and arguably resize to same host failures are less severe for holding up the release | |
| 18:46:12 | dansmith | yes, for real-world people | |
| 18:46:30 | mriedem | the ones on mtv? | |
| 18:46:31 | dansmith | realistically, the way this is going, we're going to be fixing these kinds of issues for a year | |
| 18:46:34 | dansmith | heh | |
| 18:46:40 | mriedem | no shit, this is whack a mole at it's finest | |
| 18:46:59 | dansmith | like, I'm fairly worried about how this is all going to go down | |
| 18:47:17 | cburgess | Worried about what specifically? | |
| 18:47:34 | mriedem | worried about the number of bugs that have shaken out in the last 3 weeks | |
| 18:47:35 | dansmith | cburgess: the thousand places we haven't already found and fixed | |
| 18:47:48 | cburgess | Oh.. so.. normal release cycle? :) | |
| 18:47:57 | mriedem | well, | |
| 18:48:06 | dansmith | cburgess: squared. | |
| 18:48:08 | cburgess | But seriously.. is there something that makes you more concerned this cylce then usual? | |
| 18:48:15 | cburgess | Just the volume of bugs? | |
| 18:48:23 | dansmith | cburgess: no specifically what we were just talking about | |
| 18:48:23 | mriedem | ocata was using placement but not for claims | |
| 18:48:29 | dansmith | cburgess: placement claims stuff | |
| 18:48:38 | cburgess | Ahh ok | |
| 18:48:54 | mriedem | now we're basically redoing the resource tracker with placement, but triplicating it everywhere when things fail | |
| 18:49:10 | cburgess | That means we can blame jaypipes right? | |
| 18:49:14 | mriedem | btw chet, wtf, you just show up? | |
| 18:49:18 | dansmith | doing the allocations stuff as part of RT was a huge mistake from the beginning | |
| 18:49:22 | mriedem | in our hour of need | |
| 18:49:25 | cburgess | mriedem To... ? | |
| 18:49:29 | openstackgerrit | Ildiko Vancsa proposed openstack/nova-specs master: Add spec to use cinder's new attachment API https://review.openstack.org/497552 | |
| 18:49:31 | dansmith | it should have been all clean and separate | |
| 18:49:36 | cburgess | mriedem Oh I just looked at the back scroll yeah | |
| 18:49:52 | cburgess | I try and stay current but fail mostly so when I do see something of interested I ask. | |
| 18:50:04 | edmondsw | is there any way to toggle whether an existing flavor is public or not? I don't see an update flavor API in the API docs, and the os-flavor-access APIs don't appear to do that either based on the docs... | |
| 18:50:13 | mriedem | something interesting like me and dan crying over the state of things | |
| 18:50:31 | mriedem | edmondsw: there is no update flavors api | |
| 18:50:44 | cburgess | mriedem Well.. more specifically seeing 2 cores I have a lot of respect for seriously debating concerns around stability makes me notice. | |
| 18:50:45 | mriedem | os-flavor-access restricts access per tenant | |
| 18:50:54 | mriedem | cburgess: heh as it should :) | |
| 18:50:59 | mriedem | anyway, fixing this bug quick | |
| 18:51:04 | cburgess | I want to know why so I can better steer our plans in the future of what releases to move to and when. | |
| 18:51:04 | edmondsw | mriedem right, that's what I was afraid of... boo... | |
| 18:51:16 | dansmith | cburgess: that's stage 1 concern, stage 2 is seeing us polishing up resumes | |
| 18:51:38 | dansmith | speaking of which, has anyone seen my bottle of resume polish? | |
| 18:51:44 | cburgess | dansmith Yes... very yes. | |
| 18:52:01 | cburgess | Wait there is a polish for that? Maybe thats why I've been stuck in this same job for 6 years... :P | |
| 18:52:21 | edmondsw | seems odd that we have an API to allow you to restrict access by tenant but not to open it up to all tenants | |
| 18:52:26 | dansmith | cburgess: https://cdn.dribbble.com/users/327319/screenshots/1695561/resume_polish-01_1x.png | |
| 18:52:27 | edmondsw | oh well | |
| 18:53:16 | cburgess | dansmith OMG I'm going to have to use that on social media some how. Thats brilliant. | |
| 18:58:16 | mriedem | hmm looks like remove_provider_from_instance_allocation handles the resize to same host for me | |
| 19:30:28 | openstackgerrit | Dan Smith proposed openstack/nova master: Add uuid to migration object and migrate-on-load https://review.openstack.org/496934 | |
| 19:35:27 | mriedem | ok fix incoming | |
| 19:35:27 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Cleanup allocations in failed prep_resize https://review.openstack.org/497592 | |
| 19:35:55 | mriedem | dansmith: ^ should handle both resize to same host and different hosts, but don't have a functional test for the resize to same host case | |
| 19:35:56 | mriedem | yet | |
| 19:39:45 | mriedem | oomichi: can you help review ^ too please? | |
| 19:39:52 | mriedem | we're pretty short staffed right now | |
| 19:44:05 | dansmith | mriedem: why not do the flavor->resources conversion in the caller and then just re-use the existing migration method in the RT? | |
| 19:44:41 | dansmith | other than that the method is identical to the above in terms of functionality | |
| 19:44:49 | mriedem | because the existing method uses instance.flavor | |
| 19:44:52 | mriedem | which is not the new flavor yet | |
| 19:45:25 | mriedem | the instance.flavor = new_flavor in finish_resize | |
| 19:45:29 | dansmith | ah, I misread the first line | |
| 19:45:44 | mriedem | it's sneaky | |
| 19:46:00 | dansmith | could still refactor it to take a flavor | |
| 19:46:09 | dansmith | just seems like it's too similar to duplicate | |
| 19:46:24 | mriedem | well, the name is misleading, and the error message if it fails | |
| 19:48:19 | mriedem | cdent was doing similar dedup here https://review.openstack.org/#/c/496936/ | |
| 19:48:20 | guimaluf | hi guys, I'm getting "Instance failed network setup after 1 attempt(s)" followed by "Timed out waiting for a reply to message ID Timed out waiting for a reply to message ID". My compute node run neutron-{ovs,dhcp,metadata}-agent, and neutron-server and ml2 plugin on my neutron node. Rabbitmq is working. Any clue or direction? | |