| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-07-26 | |||
| 18:48:08 | sdague | https://review.openstack.org/#/c/487478 | |
| 18:49:19 | sdague | dansmith: is there a reason you changed that? | |
| 19:01:24 | dansmith | sdague: didn't work how? but yes, we need to make that change | |
| 19:01:55 | dansmith | sdague: remember this was just a grenade hack, which meant conductor and other services were still configured to point at the cell db, not cell0 | |
| 19:02:07 | dansmith | on fresh with this they're configured wrong now | |
| 19:02:09 | bauzas | jaypipes: just reading your comment | |
| 19:02:14 | dansmith | sdague: didn't want to ask before pushing over top eh? | |
| 19:02:29 | bauzas | jaypipes: I'm a bit in and out till' 9pm my time but will think of your comment | |
| 19:02:32 | sdague | dansmith: I did, all the tests failed | |
| 19:02:36 | bauzas | oops 9.30pm | |
| 19:02:43 | sdague | cell1 db does not exist | |
| 19:03:09 | sdague | dansmith: http://logs.openstack.org/78/487478/4/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/69e2d2e/ | |
| 19:03:24 | dansmith | sdague: nova_cell1 should be created on L693 | |
| 19:03:36 | sdague | so given that jenkins -1ed that version of the patch I figured it was fine to run it the old way | |
| 19:03:49 | sdague | dansmith: well, it failed pretty hard | |
| 19:03:52 | openstackgerrit | Jan Gutter proposed openstack/nova master: Netronome SmartNIC Enablement https://review.openstack.org/483459 | |
| 19:04:27 | dansmith | it won't actually work without it, and the ironic people found it before jenkins | |
| 19:04:48 | sdague | dansmith: ok, but it doesn't pass with it | |
| 19:05:13 | dansmith | I get it, but we need to make the change | |
| 19:05:19 | sdague | http://logs.openstack.org/78/487478/4/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/69e2d2e/logs/devstacklog.txt.gz#_2017-07-26_18_12_02_728 | |
| 19:05:36 | sdague | dansmith: that's fine, you need to help me figure out what other thing is also needed for that change to work | |
| 19:05:45 | dansmith | I'm doing it now | |
| 19:07:14 | dansmith | oh it's failing on the sync before we create it.. I wonder why that's different now | |
| 19:09:33 | sdague | because this function is creating the top level nova.conf | |
| 19:10:07 | sdague | https://github.com/openstack-dev/devstack/blob/62edb2f0f64bf4ac2c75e9bbfffbad5aac5ad41c/lib/nova#L671-L680 | |
| 19:10:20 | sdague | the cell1 db is not created before the sync | |
| 19:10:24 | sdague | for the top level | |
| 19:11:01 | dansmith | right, because now that points to cell0 always in the toplevel and it's created already | |
| 19:11:16 | sdague | yes | |
| 19:12:22 | dansmith | so I switched the order of those things which I think will be fine for both cases | |
| 19:14:31 | sdague | ok, if that works, good enough | |
| 19:21:33 | sdague | dansmith: thanks for the grenade patch fix, I should put a safety wrapper around that for things that only need to set variables | |
| 19:22:03 | dansmith | sdague: np | |
| 19:33:01 | openstackgerrit | Sean Dague proposed openstack/nova master: Remove the useless fake ExtensionManager from API unittests https://review.openstack.org/486416 | |
| 19:33:11 | openstackgerrit | Sean Dague proposed openstack/nova master: Move the note about '/os-volume_boot' to the correct place https://review.openstack.org/486071 | |
| 19:37:28 | mriedem | jangutter: +2, just need jaypipes to +W now | |
| 19:38:28 | jangutter | mriedem: I sincerely owe you guys. Thanks! | |
| 19:52:19 | bauzas | jaypipes: so I'm back | |
| 19:52:28 | jaypipes | bauzas: from outer space? | |
| 19:53:22 | bauzas | jaypipes: heh, just from my backyard :p | |
| 19:54:04 | bauzas | anyway | |
| 19:54:08 | bauzas | so I saw your comment | |
| 19:54:40 | bauzas | jaypipes: (and others) https://review.openstack.org/#/c/483566/14/nova/scheduler/filter_scheduler.py@170 | |
| 19:55:34 | bauzas | so, yeah, we have a self-heal by the RT.update_available_resource(), sure, agreed | |
| 19:55:55 | bauzas | but the problem I thought was about the fact that it's only running every 60 secs by default | |
| 19:56:45 | bauzas | so we could be having a race between the time we deleted the allocations from the source host, and when the source RT is again running update_a_r() | |
| 19:57:09 | bauzas | like, we could possibly accept other instances for this host even if it's not possible | |
| 19:57:53 | bauzas | or we could possibly accept some nested RPs to be used (like SRVIO VFs) meanwhime | |
| 19:58:43 | bauzas | jaypipes: so, I was thinking, should we maybe just get the original allocations from the source host before deleting them, so that if we have problems in the scheduler, we could put them again ? | |
| 19:59:22 | bauzas | jaypipes: something like _cleanup_allocations() would return for example | |
| 19:59:58 | bauzas | that said, given how we're close to FF, I wonder if it's possible | |
| 20:00:09 | bauzas | that's it for me. | |
| 20:00:11 | bauzas | :p | |
| 20:00:21 | bauzas | roger, roger. | |
| 20:00:36 | dansmith | bauzas: see my comment in there just now | |
| 20:01:32 | dansmith | mriedem: melwitt: I'm like quad booked for shit all day.. given where we are, do we have things to discuss for cells or should we put the meeting. I think we're down to the devstack stuff and placement related bits at this point | |
| 20:01:39 | dansmith | but if we have things, we can still do it | |
| 20:02:30 | melwitt | don't think we have to have a meeting but I was just gonna say the console stuff is still up and passing jenkins. I have a devstack change up that Depends-On the stack and runs console proxies per cell and that works too | |
| 20:02:31 | bauzas | dansmith: looking | |
| 20:03:03 | dansmith | melwitt: orly, okay, I should go look at that | |
| 20:03:11 | bauzas | melwitt: dansmith: mriedem: if you need me reviewing things before pike-3 for cells v2, just lemme know | |
| 20:03:20 | dansmith | if mriedem agrees, hopefully we just had our meeting | |
| 20:03:27 | dansmith | bauzas: console stuff | |
| 20:03:32 | bauzas | ok | |
| 20:04:19 | melwitt | CONSOLEZ. I'm about to update it to redact tokens in the debug logs but other than that, it's been ready. mostly unchanged from when PaulMurray was working on it | |
| 20:04:49 | mriedem | ok to ditch the meeting | |
| 20:05:09 | mriedem | what do we lose if we don't have the console proxy stuff done? | |
| 20:05:27 | mriedem | because < 24 hours is tough for something that hasn't had review yet | |
| 20:07:12 | melwitt | mriedem: I think just inability to shard console proxies. current state is the token auth cache and proxies are global. the console stuff moves token auth storage to the cell databases and then shards proxies, one per cell | |
| 20:07:45 | melwitt | and starts on the deprecation timer on eliminating the consoleauth service | |
| 20:09:12 | bauzas | dansmith: looks good for me with the plan | |
| 20:09:14 | dansmith | sounds like a good candidate to punt then | |
| 20:09:19 | dansmith | bauzas: ack | |
| 20:09:51 | bauzas | dansmith: I was just thinking of restoring the original allocation but honestly having both allocations for the source and destination hosts make more sense in terms of "resource usage" | |
| 20:10:10 | bauzas | because when you wanna move, you need to make sure you have double room | |
| 20:10:15 | bauzas | until the move is done | |
| 20:10:19 | melwitt | dansmith: yeah. I think the main concern I had was if a change in deployment topology being out-of-sync with the rest of the multicell changes, but I think based on our current state, superconductor isn't going to be a thing yet, right? | |
| 20:11:29 | dansmith | melwitt: it is, merged in devstack now.. not sure I get the relation | |
| 20:11:42 | dansmith | melwitt: or you mean a change in deployment from pike->queens? | |
| 20:11:45 | melwitt | or rather, I don't know what the current state of multicell is in regard to what we will document for users. I think I asked the question on the etherpad, what does multi-tier mean vs multi-cell? | |
| 20:12:26 | dansmith | no difference, I just used the term once to mean something specific and someone started saying it I think | |
| 20:12:26 | melwitt | like, when we will have the communication of "you'll need to change your deployment of services like this" I was thinking it might be less confusing if the console proxy run location changes coincided with that | |
| 20:12:59 | dansmith | the only people affected by such a change would be people that would split out a cell in pike, and then upgrade to queens I think | |
| 20:13:17 | melwitt | instead of letting someone get started with multicell and global console proxies and then in another release saying, "oh yeah, go back and change what you already deployed" | |
| 20:13:19 | dansmith | and anyone that does that won't have reschedules and affinity checks in pike anyway | |
| 20:13:34 | dansmith | yeah, so that'd be the only concern I think | |
| 20:13:37 | melwitt | right | |
| 20:13:40 | dansmith | I think that the migration wouldn't be bad though, | |
| 20:14:03 | dansmith | because you can just start up console in the cell prior to the upgrade and it won't do anything until after things roll | |
| 20:14:13 | dansmith | up to mriedem to decide which he thinks is less problematic/risky | |
| 20:14:25 | dansmith | merging early, or documenting that change to the people that actually split a cell in pike | |
| 20:14:31 | melwitt | yeah, I guess not. it would be like, stand up console proxies per cell and have to leave the global ones running too until no outstanding tokens are using them anymore | |
| 20:14:45 | dansmith | well, as I've said, | |
| 20:14:52 | dansmith | I don't think that resetting tokens is a big deal | |
| 20:15:04 | dansmith | you just don't want to have a time where you can't get new tokens | |
| 20:15:20 | openstackgerrit | Eric Fried proposed openstack/nova master: nova.utils.get_service_url() https://review.openstack.org/458257 | |
| 20:15:22 | melwitt | yeah, true | |
| 20:15:43 | efried | mriedem jaypipes johnthetubaguy (mordred) ^ | |
| 20:16:29 | mordred | efried: woot! | |
| 20:16:47 | efried | May need a little help making sure I find all the right places that need to be touched for the followup (other places in nova that talk to endpoints) | |