Earlier  
Posted Nick Remark
#openstack-nova - 2017-07-26
19:37:28 mriedem jangutter: +2, just need jaypipes to +W now
19:38:28 jangutter mriedem: I sincerely owe you guys. Thanks!
19:52:19 bauzas jaypipes: so I'm back
19:52:28 jaypipes bauzas: from outer space?
19:53:22 bauzas jaypipes: heh, just from my backyard :p
19:54:04 bauzas anyway
19:54:08 bauzas so I saw your comment
19:54:40 bauzas jaypipes: (and others) https://review.openstack.org/#/c/483566/14/nova/scheduler/filter_scheduler.py@170
19:55:34 bauzas so, yeah, we have a self-heal by the RT.update_available_resource(), sure, agreed
19:55:55 bauzas but the problem I thought was about the fact that it's only running every 60 secs by default
19:56:45 bauzas so we could be having a race between the time we deleted the allocations from the source host, and when the source RT is again running update_a_r()
19:57:09 bauzas like, we could possibly accept other instances for this host even if it's not possible
19:57:53 bauzas or we could possibly accept some nested RPs to be used (like SRVIO VFs) meanwhime
19:58:43 bauzas jaypipes: so, I was thinking, should we maybe just get the original allocations from the source host before deleting them, so that if we have problems in the scheduler, we could put them again ?
19:59:22 bauzas jaypipes: something like _cleanup_allocations() would return for example
19:59:58 bauzas that said, given how we're close to FF, I wonder if it's possible
20:00:09 bauzas that's it for me.
20:00:11 bauzas :p
20:00:21 bauzas roger, roger.
20:00:36 dansmith bauzas: see my comment in there just now
20:01:32 dansmith mriedem: melwitt: I'm like quad booked for shit all day.. given where we are, do we have things to discuss for cells or should we put the meeting. I think we're down to the devstack stuff and placement related bits at this point
20:01:39 dansmith but if we have things, we can still do it
20:02:30 melwitt don't think we have to have a meeting but I was just gonna say the console stuff is still up and passing jenkins. I have a devstack change up that Depends-On the stack and runs console proxies per cell and that works too
20:02:31 bauzas dansmith: looking
20:03:03 dansmith melwitt: orly, okay, I should go look at that
20:03:11 bauzas melwitt: dansmith: mriedem: if you need me reviewing things before pike-3 for cells v2, just lemme know
20:03:20 dansmith if mriedem agrees, hopefully we just had our meeting
20:03:27 dansmith bauzas: console stuff
20:03:32 bauzas ok
20:04:19 melwitt CONSOLEZ. I'm about to update it to redact tokens in the debug logs but other than that, it's been ready. mostly unchanged from when PaulMurray was working on it
20:04:49 mriedem ok to ditch the meeting
20:05:09 mriedem what do we lose if we don't have the console proxy stuff done?
20:05:27 mriedem because < 24 hours is tough for something that hasn't had review yet
20:07:12 melwitt mriedem: I think just inability to shard console proxies. current state is the token auth cache and proxies are global. the console stuff moves token auth storage to the cell databases and then shards proxies, one per cell
20:07:45 melwitt and starts on the deprecation timer on eliminating the consoleauth service
20:09:12 bauzas dansmith: looks good for me with the plan
20:09:14 dansmith sounds like a good candidate to punt then
20:09:19 dansmith bauzas: ack
20:09:51 bauzas dansmith: I was just thinking of restoring the original allocation but honestly having both allocations for the source and destination hosts make more sense in terms of "resource usage"
20:10:10 bauzas because when you wanna move, you need to make sure you have double room
20:10:15 bauzas until the move is done
20:10:19 melwitt dansmith: yeah. I think the main concern I had was if a change in deployment topology being out-of-sync with the rest of the multicell changes, but I think based on our current state, superconductor isn't going to be a thing yet, right?
20:11:29 dansmith melwitt: it is, merged in devstack now.. not sure I get the relation
20:11:42 dansmith melwitt: or you mean a change in deployment from pike->queens?
20:11:45 melwitt or rather, I don't know what the current state of multicell is in regard to what we will document for users. I think I asked the question on the etherpad, what does multi-tier mean vs multi-cell?
20:12:26 melwitt like, when we will have the communication of "you'll need to change your deployment of services like this" I was thinking it might be less confusing if the console proxy run location changes coincided with that
20:12:26 dansmith no difference, I just used the term once to mean something specific and someone started saying it I think
20:12:59 dansmith the only people affected by such a change would be people that would split out a cell in pike, and then upgrade to queens I think
20:13:17 melwitt instead of letting someone get started with multicell and global console proxies and then in another release saying, "oh yeah, go back and change what you already deployed"
20:13:19 dansmith and anyone that does that won't have reschedules and affinity checks in pike anyway
20:13:34 dansmith yeah, so that'd be the only concern I think
20:13:37 melwitt right
20:13:40 dansmith I think that the migration wouldn't be bad though,
20:14:03 dansmith because you can just start up console in the cell prior to the upgrade and it won't do anything until after things roll
20:14:13 dansmith up to mriedem to decide which he thinks is less problematic/risky
20:14:25 dansmith merging early, or documenting that change to the people that actually split a cell in pike
20:14:31 melwitt yeah, I guess not. it would be like, stand up console proxies per cell and have to leave the global ones running too until no outstanding tokens are using them anymore
20:14:45 dansmith well, as I've said,
20:14:52 dansmith I don't think that resetting tokens is a big deal
20:15:04 dansmith you just don't want to have a time where you can't get new tokens
20:15:20 openstackgerrit Eric Fried proposed openstack/nova master: nova.utils.get_service_url() https://review.openstack.org/458257
20:15:22 melwitt yeah, true
20:15:43 efried mriedem jaypipes johnthetubaguy (mordred) ^
20:16:29 mordred efried: woot!
20:16:47 efried May need a little help making sure I find all the right places that need to be touched for the followup (other places in nova that talk to endpoints)
20:17:27 bauzas mriedem: jangutter: I have a concern with https://review.openstack.org/#/c/483459/17
20:17:31 efried mordred The difference in ksa usage is dissapointingly tiny in proportion to the amount of ksa work done :)
20:17:42 efried mordred But os-service-types is gold
20:17:46 bauzas mriedem: jangutter: AFAICT, it's depending on an os-vif change, right?
20:17:56 mordred efried: you're setting ks_adapter.interface to a list in utils - isn't there some way to set a conf default or something?
20:17:57 mriedem bauzas: which is already merged and released
20:17:58 mordred efried: woot!
20:18:07 jangutter bauzas: correct
20:18:09 bauzas mriedem: oh nevermind, just saw it
20:18:15 efried mordred The conf options come directly from ksa.
20:18:21 bauzas mriedem: I thought we didn't released it
20:18:32 mordred yah- I thought there was maybe some way to set/override the deafult value for one of them or something
20:19:01 efried mordred Mm, not sure about that. Doug would probably know.
20:19:14 jangutter bauzas: yeah, it snuck in right as the door closed.
20:19:21 efried mordred Not completely sure we would want to do that, exactly.
20:19:35 mordred efried: nod
20:19:38 bauzas jangutter: happy for you
20:19:45 efried I guess I can see the advantage for automatic conf generation or something.
20:20:01 openstackgerrit melanie witt proposed openstack/nova master: Add periodic task to clean expired console tokens https://review.openstack.org/325381
20:20:02 openstackgerrit melanie witt proposed openstack/nova master: Add console connection object https://review.openstack.org/320063
20:20:02 openstackgerrit melanie witt proposed openstack/nova master: Use ConsoleConnection object to generate authorizations https://review.openstack.org/325414
20:20:03 openstackgerrit melanie witt proposed openstack/nova master: Convert websocketproxy to use db for token validation https://review.openstack.org/333990
20:20:03 openstackgerrit melanie witt proposed openstack/nova master: Add access_url_base to console_auth_tokens table https://review.openstack.org/334614
20:20:04 openstackgerrit melanie witt proposed openstack/nova master: Add console_auth_token_get() method to DB API https://review.openstack.org/481700
20:20:54 jangutter bauzas: thanks, I owe the gang some serious drinks.
20:21:43 mriedem melwitt: so i think to summarize the statement for multi-cell in pike, is cern will be ok with it since they don't do build retries anyway, but if you rely on retries then you shouldn't do multi-cell in pike, since we don't have the alternatives stuff done
20:22:09 mriedem or if you rely on server group affinity/anti-affinity since the computes and cell conductors can't hit the scheduler or api db
20:22:15 mriedem dansmith: ^ keep me honest on that 2nd one
20:23:29 melwitt yeah, that makes sense. I just wasn't sure what the difference between multi-tier and multi-cell was but if they're synonyms then I think I understand the state
20:23:47 openstackgerrit Merged openstack/python-novaclient master: Change Service repr to use self.id always https://review.openstack.org/487502
20:23:51 bauzas mriedem: dansmith: melwitt: correct me if I'm wrong but you can multi-cell and retry if you are able to upcall ?
20:24:12 bauzas the only problem is when you usually isolate your MQs
20:25:25 mriedem melwitt: multi-tier to me is superconductor, and it's not really worth doing multi-cell if you can't do multi-tier with superconductor, because then your cells are all going to be blasting to every other cell conductor - which breaks a bunch of stuff
20:25:32 mriedem dan mentioned this on a hangout yesterday with me and jay
20:26:09 mriedem well there is no cell conductor,

Earlier   Later