Earlier  
Posted Nick Remark
#openstack-nova - 2017-07-26
17:03:18 cfriesen mriedem: no need for them...the computes don't talk to anything outside the cluster and they know the hostname of everything in the cluster.
17:05:06 dansmith cfriesen: except you just said the reason to have them :)
17:05:31 jangutter mriedem, jaypipes: sean-k-mooney gave his +1 on https://review.openstack.org/#/c/483459 have I addressed your concerns too?
17:05:49 mriedem jangutter: i'll look again after lunch
17:05:50 cfriesen dansmith: heh...I've proposed https://review.openstack.org/#/c/487494 which is the proper fix. (Once I add in the error for hostname not found.)
17:06:09 jangutter mriedem: thanks!
17:06:57 dansmith cfriesen: except that breaks the current, legit behavior
17:07:21 cfriesen dansmith: we agreed at the last PTG that the current pattern-match behaviour didn't make sense and should be changed.
17:07:31 dansmith we did?
17:07:34 cfriesen yep
17:08:11 dansmith if I have compute0001.cellN.foo.com in each cell, I can evacuate them in parallel with no shared infrastructure between them and have no problems
17:08:43 cfriesen dansmith: the help text for the cases in question are written as affecting a single host, but they use a pattern match which could affect multiple hosts accidentally
17:08:45 dansmith I can't imagine why we wouldn't want a --strict flag to enable that new behavior
17:09:18 dansmith but the behavior wins over incorrect docs right?
17:09:25 cburgess dansmith I'm not following. I actually didn't know it pattern matches and I'm having a hard time figuring out why you would want that. Can you explain the use case more?
17:09:52 cfriesen dansmith: from my notes it was on the Friday session at the PTG
17:10:06 dansmith cburgess: lets say I'm marching through compute nodes to do upgrades, getting all the instances off each one you're going to upgrade first
17:10:15 cfriesen and I think it's counterintuitive that a command called "nova host-evacuate" would affect multiple hosts
17:10:45 dansmith cburgess: and you want to do batches, either in parallel across cells or just the first N nodes at a time
17:11:13 dansmith cfriesen: I'm not saying it wouldn't have made sense to make it behave that way originally, I just don't see why we should change it instead of offering the option for --strict
17:11:14 cburgess dansmith And the argument is the client should enable that rather then the admin just issuing N calls?
17:11:26 dansmith cburgess: this is admin only
17:11:33 cfriesen dansmith: if you specify "compute-1 it'll affect compute-1, compute-10 to compute-19, compute-100 to compute-199, etc.
17:11:50 dansmith cburgess: you meant the CLI I guess. I'm just saying it _does_
17:12:11 dansmith cfriesen: yeah, I understand how pattern matching works.. hence my example above with not insane compute node hostnames :)
17:12:42 cburgess dansmith I'm with cfriesen on this in that its its extremely counter-intuitive to me. But I get what you are saying about its been that way for a long time so we probably need to be careful changing it now and possible use a flag to do that.
17:13:19 dansmith cburgess: I'm with you too in saying that it probably shouldn't have been done this way in the beginning,
17:13:36 dansmith but it clearly intended to pattern match, so I'm guessing the intent was, you know, to be able to do that :)
17:14:07 cfriesen I'm not sure it was intentional...I think it was just fallout of the fact that we don't have a way to look up a single hypervisor by name
17:14:09 dansmith and as we always say, the api docs don't matter, the actual behavior is what matters and what people will build dependencies on
17:14:10 cburgess dansmith Granted... we call the API directly for this rather then use the CLI because there are other issues around limiting the number of in-flight migrations etc. So this is partially just me saying "That don't make no sense".
17:14:43 cburgess dansmith I get it... like I said, mostly be saying "That don't make no sense." But hey... its what we ship.
17:14:49 dansmith cburgess: not so much across hosts, which is where the matching is here, and we added the live migration throttle, but yeah agreed
17:15:04 cburgess dansmith live migration throttle?
17:15:20 dansmith cfriesen: if it wasn't intentional, it would have failed if the match was zero or >1, or just taken the first match, but it clearly adds all the resulting hosts, so .. clearly intentional, IMHO
17:15:21 cfriesen if we're going to change the docs per mriedem's patch...is it worth me updating to include the --strict hostname matching mode or just leave it as is?
17:15:32 dansmith cburgess: yeah you can limit how many migrations in parallel a single host will make
17:16:02 cburgess dansmith This is a nova.conf option or API thing... or CLI thing? Sorry a bit slow this morning it seems.
17:16:03 dansmith cfriesen: I'm happy with your patch if you make it only do that thing on --strict
17:16:17 dansmith cfriesen: or if you want to go down the path of requiring either --strict or --loose and do the major version bump dance
17:16:25 dansmith cburgess: yeah, nova-compute conf
17:16:27 dansmith cburgess: hold on
17:16:55 dansmith cburgess: https://github.com/openstack/nova/blob/master/nova/conf/compute.py#L588-L601
17:16:58 dansmith cburgess: you're welcome :)
17:17:25 dansmith cburgess: we have a build limit too, in case you hadn't seen it
17:17:35 dansmith cburgess: to avoid building 20 instances in parallel on a single compute node
17:17:39 dansmith s/building/failing to build/ :)
17:18:34 cburgess dansmith Interesting... thanks. I will have to read up on this. I vaguely recall some conversation about this several summits ago.
17:21:14 dansmith I feel like this is a 3-coffee day
17:30:07 cburgess dansmith So if I'm reading the code right... the max_concurrent_live_migrations option protects the source hypervisor from having more then 1 migrations at a time but the destination isn't protected at all. Does this jive with your understanding?
17:30:38 dansmith cburgess: yep, it's outbound
17:30:55 dansmith cburgess: could do the same for inbound, or obey the build counter for inbounds
17:31:17 dansmith the reason was,
17:31:40 dansmith on host-evac-live, you're necessarily hitting one compute node for outbound, but scheduler could spread the targets out
17:31:52 dansmith obviously if you're packing and have an empty host, they'll all want to go to the same place
17:32:32 cburgess dansmith inbound seems like it would be harder... you would have to modify the flow so that the source makes a call to the dest to try and acquire a migration slot or something like that and then keep retrying to get it before it does the migration. Seems like it would be tough to do right an not end up with deadlocks.
17:33:04 openstackgerrit Merged openstack/python-novaclient master: Microversion 2.53 - services and hypervisors using UUIDs https://review.openstack.org/485435
17:33:09 dansmith cburgess: ah, yeah I guess because we make a blocking call for live migrate but not for build, fair point
17:34:20 cburgess dansmith Yeah... it would be very tricky.. which is a bummer since protecting the dest is something we would want to do as well (we actually try very hard to ensure we never have more then 1 in-coming migration at time in liberty due to some pretty nasty os-brick race conditions).
17:35:40 cburgess dansmith But its good to know about these 2 options. I think we will tune the max_concurrent_builds down some to help with performance.
17:36:06 dansmith cburgess: donuts appreciated :P
17:36:53 cburgess dansmith hehe... yeah I didn't do donuts in Boston because mriedem said we wouldn't really have the right setup. I'll look into the donuts making a return in Denver.
17:37:24 dansmith cburgess: I'm only joking.. you're way over your donuts requirement quota
17:38:28 vdrok dansmith: ouch http://logs.openstack.org/58/487458/2/check/gate-tempest-dsvm-ironic-pxe_ipmitool-postgres-ubuntu-xenial-nv/40fb4fb/logs/devstacklog.txt.gz waiting for hypervisors fail
17:39:05 openstackgerrit OpenStack Proposal Bot proposed openstack/nova master: Updated from global requirements https://review.openstack.org/487473
17:40:27 vdrok the variable picked up by nova http://logs.openstack.org/58/487458/2/check/gate-tempest-dsvm-ironic-pxe_ipmitool-postgres-ubuntu-xenial-nv/40fb4fb/logs/devstacklog.txt.gz#_2017-07-26_16_06_41_216
17:40:46 dansmith vdrok: gdi, nothing but bad news from you ever
17:40:56 dansmith vdrok: you never just stop by to say hi, always to say someting is broken! :)
17:40:56 vdrok :D
17:41:20 vdrok I promise to start the next day with 'good morning' to nova channel :)
17:41:25 cburgess dansmith You have sent me down a rabbit whole here reading all the options with that link to the code. Its nice to have all the options in one place but man its a bit scary reading some of this stuff.
17:42:05 mriedem is there anything i need to care about in the scrollback?
17:42:11 mriedem besides chet letting us down on donuts in boston?
17:42:14 cburgess mriedem donuts
17:42:18 cburgess lol
17:42:21 dansmith vdrok: http://logs.openstack.org/58/487458/2/check/gate-tempest-dsvm-ironic-pxe_ipmitool-postgres-ubuntu-xenial-nv/40fb4fb/logs/screen-n-cpu.txt.gz#_Jul_26_16_08_05_910044
17:42:28 dansmith vdrok: looks like it created the compute node...
17:43:22 vdrok yup, but for some reason it's not picked up by hypervisor-stats
17:44:47 dansmith vdrok: and never by discover either
17:47:46 dansmith vdrok: okay I think this might be because this hack was for grenade and it's not complete enough for a fresh install
17:48:06 dansmith vdrok: I think we're pointing at super-conductor (which we want) but it's configured to point at cell0 (which we do not)
17:48:14 dansmith vdrok: so give me a few to mull this over
17:49:31 openstackgerrit Jackie Truong proposed openstack/nova master: Add trusted certificates to InstanceExtras https://review.openstack.org/457711
17:56:29 dansmith vdrok: updated the devstack change
17:57:51 vdrok dansmith: thanks! will come back if it fails again :D
17:57:55 vdrok good night
17:58:07 dansmith vdrok: o/
18:30:22 mriedem jangutter: really only nits in the release note now https://review.openstack.org/#/c/483459/
18:30:28 mriedem but see what you think
18:32:51 jangutter mriedem: I'm working on getting our docs cleared for general release on our support site. Currently customers should have subscriptions, but they should be generally available by the time Pike gets released.
18:33:17 jangutter mriedem: should I respin and remove the soon?
18:47:58 jaypipes bauzas: answered. sorry, went to lunch
18:47:58 sdague dansmith: your cell1 db change didn't work
18:48:08 sdague https://review.openstack.org/#/c/487478
18:49:19 sdague dansmith: is there a reason you changed that?
19:01:24 dansmith sdague: didn't work how? but yes, we need to make that change
19:01:55 dansmith sdague: remember this was just a grenade hack, which meant conductor and other services were still configured to point at the cell db, not cell0
19:02:07 dansmith on fresh with this they're configured wrong now
19:02:09 bauzas jaypipes: just reading your comment
19:02:14 dansmith sdague: didn't want to ask before pushing over top eh?

Earlier   Later