| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-07-26 | |||
| 16:18:05 | bauzas | jaypipes: I know we do this but on a periodic base | |
| 16:19:18 | bauzas | jaypipes: but my question is more about a possible race condition between the time we delete the allocations and the source compute runs again the RT that will self-heal the allocations | |
| 16:20:47 | jaypipes | bauzas: on call | |
| 16:20:52 | jaypipes | gimme few | |
| 16:21:15 | bauzas | mriedem: +2d | |
| 16:21:21 | bauzas | mriedem: just find another peep :p | |
| 16:21:46 | bauzas | oh oops | |
| 16:21:50 | bauzas | s/peep/peer | |
| 16:24:46 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: Enable cold migration with target host(1/2) https://review.openstack.org/408955 | |
| 16:32:07 | bauzas | jaypipes: need to go awol for dinner, but you can ping me later or drop a note if you think I scared for nothing | |
| 16:32:18 | jaypipes | bauzas: will do! | |
| 16:36:21 | openstackgerrit | Chris Dent proposed openstack/nova master: Optional separate database for placement API https://review.openstack.org/362766 | |
| 16:47:24 | openstackgerrit | Matt Riedemann proposed openstack/python-novaclient master: Be clear about hypevisors.search used in a few CLIs https://review.openstack.org/487513 | |
| 16:47:28 | mriedem | cfriesen: dansmith: ^ | |
| 16:47:32 | mriedem | that's backportable to stable | |
| 16:47:40 | mriedem | changing the behavior of those CLIs is not | |
| 16:49:14 | openstackgerrit | Matthew Booth proposed openstack/nova master: Fix scope of errors_out_migration in resize_instance https://review.openstack.org/487495 | |
| 16:49:15 | openstackgerrit | Matthew Booth proposed openstack/nova master: Split Compute.errors_out_migration into a separate contextmanager https://review.openstack.org/485734 | |
| 16:49:15 | openstackgerrit | Matthew Booth proposed openstack/nova master: Ensure errors_out_migration errors out migration https://review.openstack.org/479802 | |
| 16:49:16 | openstackgerrit | Matthew Booth proposed openstack/nova master: Fix scope of errors_out_migration in finish_resize https://review.openstack.org/487515 | |
| 16:51:46 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: Enable cold migration with target host(2/2) https://review.openstack.org/408964 | |
| 16:51:51 | mriedem | sdague: grenade exploded http://logs.openstack.org/85/487485/1/check/gate-grenade-dsvm-neutron-ubuntu-xenial/7ab03fb/logs/grenade.sh.txt.gz#_2017-07-26_15_59_28_974 | |
| 16:52:22 | mriedem | but, that's appropriate for a project called 'grenade' i guess | |
| 16:53:12 | cfriesen | mriedem: who approved "host-evacuate-live" anyway? :) | |
| 16:56:38 | cfriesen | mriedem: the suggestion to use an fqdn assumes that you have fqdns in your cluster | |
| 16:56:54 | mdbooth | cfriesen: That tool's old as the hills, isn't it? | |
| 16:58:31 | cfriesen | in ours we just have hostnames "compute-1, compute-10, compute-100, etc" so there is no way to run "host-evacuate" on just compute-1 without changing the novaclient code. | |
| 16:58:59 | cfriesen | mdbooth: looks like, yes. really "evacuate" should have been "resurrect" :) | |
| 16:59:11 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: Enable cold migration with target host(2/2) https://review.openstack.org/408964 | |
| 17:00:09 | mdbooth | cfriesen: I wrote this a while back, btw: https://gist.github.com/mdbooth/163f5fdf47ab45d7addd | |
| 17:00:26 | mdbooth | No idea if it's useful to you, or it still works for that matter :) | |
| 17:00:32 | mdbooth | Although I'd hope the latter | |
| 17:01:04 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: api-ref: Add parameters in cold migrate action https://review.openstack.org/410042 | |
| 17:01:53 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: List/show all server migration types (1/2) https://review.openstack.org/430608 | |
| 17:01:54 | cfriesen | mdbooth: we've got our own management component that does something similar. | |
| 17:01:59 | mriedem | cfriesen: why wouldn't you have fqdns on the compute hosts? | |
| 17:02:09 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: List/show all server migration types (2/2) https://review.openstack.org/459483 | |
| 17:02:22 | mdbooth | cfriesen: Not surprised. Seems like a pretty common thing to want. | |
| 17:03:18 | cfriesen | mriedem: no need for them...the computes don't talk to anything outside the cluster and they know the hostname of everything in the cluster. | |
| 17:05:06 | dansmith | cfriesen: except you just said the reason to have them :) | |
| 17:05:31 | jangutter | mriedem, jaypipes: sean-k-mooney gave his +1 on https://review.openstack.org/#/c/483459 have I addressed your concerns too? | |
| 17:05:49 | mriedem | jangutter: i'll look again after lunch | |
| 17:05:50 | cfriesen | dansmith: heh...I've proposed https://review.openstack.org/#/c/487494 which is the proper fix. (Once I add in the error for hostname not found.) | |
| 17:06:09 | jangutter | mriedem: thanks! | |
| 17:06:57 | dansmith | cfriesen: except that breaks the current, legit behavior | |
| 17:07:21 | cfriesen | dansmith: we agreed at the last PTG that the current pattern-match behaviour didn't make sense and should be changed. | |
| 17:07:31 | dansmith | we did? | |
| 17:07:34 | cfriesen | yep | |
| 17:08:11 | dansmith | if I have compute0001.cellN.foo.com in each cell, I can evacuate them in parallel with no shared infrastructure between them and have no problems | |
| 17:08:43 | cfriesen | dansmith: the help text for the cases in question are written as affecting a single host, but they use a pattern match which could affect multiple hosts accidentally | |
| 17:08:45 | dansmith | I can't imagine why we wouldn't want a --strict flag to enable that new behavior | |
| 17:09:18 | dansmith | but the behavior wins over incorrect docs right? | |
| 17:09:25 | cburgess | dansmith I'm not following. I actually didn't know it pattern matches and I'm having a hard time figuring out why you would want that. Can you explain the use case more? | |
| 17:09:52 | cfriesen | dansmith: from my notes it was on the Friday session at the PTG | |
| 17:10:06 | dansmith | cburgess: lets say I'm marching through compute nodes to do upgrades, getting all the instances off each one you're going to upgrade first | |
| 17:10:15 | cfriesen | and I think it's counterintuitive that a command called "nova host-evacuate" would affect multiple hosts | |
| 17:10:45 | dansmith | cburgess: and you want to do batches, either in parallel across cells or just the first N nodes at a time | |
| 17:11:13 | dansmith | cfriesen: I'm not saying it wouldn't have made sense to make it behave that way originally, I just don't see why we should change it instead of offering the option for --strict | |
| 17:11:14 | cburgess | dansmith And the argument is the client should enable that rather then the admin just issuing N calls? | |
| 17:11:26 | dansmith | cburgess: this is admin only | |
| 17:11:33 | cfriesen | dansmith: if you specify "compute-1 it'll affect compute-1, compute-10 to compute-19, compute-100 to compute-199, etc. | |
| 17:11:50 | dansmith | cburgess: you meant the CLI I guess. I'm just saying it _does_ | |
| 17:12:11 | dansmith | cfriesen: yeah, I understand how pattern matching works.. hence my example above with not insane compute node hostnames :) | |
| 17:12:42 | cburgess | dansmith I'm with cfriesen on this in that its its extremely counter-intuitive to me. But I get what you are saying about its been that way for a long time so we probably need to be careful changing it now and possible use a flag to do that. | |
| 17:13:19 | dansmith | cburgess: I'm with you too in saying that it probably shouldn't have been done this way in the beginning, | |
| 17:13:36 | dansmith | but it clearly intended to pattern match, so I'm guessing the intent was, you know, to be able to do that :) | |
| 17:14:07 | cfriesen | I'm not sure it was intentional...I think it was just fallout of the fact that we don't have a way to look up a single hypervisor by name | |
| 17:14:09 | dansmith | and as we always say, the api docs don't matter, the actual behavior is what matters and what people will build dependencies on | |
| 17:14:10 | cburgess | dansmith Granted... we call the API directly for this rather then use the CLI because there are other issues around limiting the number of in-flight migrations etc. So this is partially just me saying "That don't make no sense". | |
| 17:14:43 | cburgess | dansmith I get it... like I said, mostly be saying "That don't make no sense." But hey... its what we ship. | |
| 17:14:49 | dansmith | cburgess: not so much across hosts, which is where the matching is here, and we added the live migration throttle, but yeah agreed | |
| 17:15:04 | cburgess | dansmith live migration throttle? | |
| 17:15:20 | dansmith | cfriesen: if it wasn't intentional, it would have failed if the match was zero or >1, or just taken the first match, but it clearly adds all the resulting hosts, so .. clearly intentional, IMHO | |
| 17:15:21 | cfriesen | if we're going to change the docs per mriedem's patch...is it worth me updating to include the --strict hostname matching mode or just leave it as is? | |
| 17:15:32 | dansmith | cburgess: yeah you can limit how many migrations in parallel a single host will make | |
| 17:16:02 | cburgess | dansmith This is a nova.conf option or API thing... or CLI thing? Sorry a bit slow this morning it seems. | |
| 17:16:03 | dansmith | cfriesen: I'm happy with your patch if you make it only do that thing on --strict | |
| 17:16:17 | dansmith | cfriesen: or if you want to go down the path of requiring either --strict or --loose and do the major version bump dance | |
| 17:16:25 | dansmith | cburgess: yeah, nova-compute conf | |
| 17:16:27 | dansmith | cburgess: hold on | |
| 17:16:55 | dansmith | cburgess: https://github.com/openstack/nova/blob/master/nova/conf/compute.py#L588-L601 | |
| 17:16:58 | dansmith | cburgess: you're welcome :) | |
| 17:17:25 | dansmith | cburgess: we have a build limit too, in case you hadn't seen it | |
| 17:17:35 | dansmith | cburgess: to avoid building 20 instances in parallel on a single compute node | |
| 17:17:39 | dansmith | s/building/failing to build/ :) | |
| 17:18:34 | cburgess | dansmith Interesting... thanks. I will have to read up on this. I vaguely recall some conversation about this several summits ago. | |
| 17:21:14 | dansmith | I feel like this is a 3-coffee day | |
| 17:30:07 | cburgess | dansmith So if I'm reading the code right... the max_concurrent_live_migrations option protects the source hypervisor from having more then 1 migrations at a time but the destination isn't protected at all. Does this jive with your understanding? | |
| 17:30:38 | dansmith | cburgess: yep, it's outbound | |
| 17:30:55 | dansmith | cburgess: could do the same for inbound, or obey the build counter for inbounds | |
| 17:31:17 | dansmith | the reason was, | |
| 17:31:40 | dansmith | on host-evac-live, you're necessarily hitting one compute node for outbound, but scheduler could spread the targets out | |
| 17:31:52 | dansmith | obviously if you're packing and have an empty host, they'll all want to go to the same place | |
| 17:32:32 | cburgess | dansmith inbound seems like it would be harder... you would have to modify the flow so that the source makes a call to the dest to try and acquire a migration slot or something like that and then keep retrying to get it before it does the migration. Seems like it would be tough to do right an not end up with deadlocks. | |
| 17:33:04 | openstackgerrit | Merged openstack/python-novaclient master: Microversion 2.53 - services and hypervisors using UUIDs https://review.openstack.org/485435 | |
| 17:33:09 | dansmith | cburgess: ah, yeah I guess because we make a blocking call for live migrate but not for build, fair point | |
| 17:34:20 | cburgess | dansmith Yeah... it would be very tricky.. which is a bummer since protecting the dest is something we would want to do as well (we actually try very hard to ensure we never have more then 1 in-coming migration at time in liberty due to some pretty nasty os-brick race conditions). | |
| 17:35:40 | cburgess | dansmith But its good to know about these 2 options. I think we will tune the max_concurrent_builds down some to help with performance. | |
| 17:36:06 | dansmith | cburgess: donuts appreciated :P | |
| 17:36:53 | cburgess | dansmith hehe... yeah I didn't do donuts in Boston because mriedem said we wouldn't really have the right setup. I'll look into the donuts making a return in Denver. | |