Earlier  
Posted Nick Remark
#openstack-nova - 2021-01-05
14:59:43 bauzas it was
14:59:45 sean-k-mooney dansmith: so just form a security and manatiance point of view we would have needed to eventually move
14:59:55 dansmith bauzas: roger
15:00:02 dansmith sean-k-mooney: no, I know ;)
15:00:03 bauzas and I've been told intermediate versions were providing both UIs
15:00:24 bauzas but we were so lagging that when we jumped straight, gerrit removed the old UI meanwhile
15:00:24 sean-k-mooney yes
15:00:30 sean-k-mooney they did that for about a year or so
15:00:34 bauzas \o/
15:01:13 sean-k-mooney yep thats effectivly what happened
15:09:37 bauzas this would make our conversations much simplier
15:09:55 stephenfin or, you know, change your nick :)
15:10:25 bauzas you can't imagine how many colleagues were thinking that my last name was ending with an 's'
15:10:28 bauzas damn IRC
15:12:36 bauzas gosh, the cyborg shelve patch is not exactly hairy, but I'd have preferred it being split between the API change and the RPC changes
15:13:55 stephenfin from my brief look, that's probably wise
15:14:00 stephenfin API last, of course
15:14:02 bauzas gibi: any idea why there are conductor changes with https://review.opendev.org/c/openstack/nova/+/729563/26/nova/conductor/manager.py ?
15:14:47 bauzas stephenfin: from someone who fixed some RPC compat break from the last cyborg patch, please understand my cautiousness
15:15:32 gibi bauzas: unshelve going through the conductor as it needs to call the scheduler
15:15:35 gibi after shelve offload
15:16:30 bauzas gibi: in rebuild_instance() ?
15:17:07 bauzas anyway, taxi time
15:17:18 bauzas will figure this out when I'm back, 15 min-ish
15:17:19 sean-k-mooney rebuild hits the schduler too to assert the new image is valid for the current host so maybe that
15:17:34 gibi bauzas: I have to guess it is historical, this patch went thorough many many revision
15:17:43 bauzas sean-k-mooney: I just honestly feel they added some fix in the same change
15:17:44 gibi I will find the reason
15:17:49 bauzas but that looks extra
15:18:35 sean-k-mooney ill try to take a look eairlier today although i need to start working on something else too
15:18:43 gibi bauzas: one thing that _create_and_bind_arq_for_instance() has been changed and that is called from multiple places
15:19:45 sean-k-mooney they are changin form host to host.nodename
15:20:23 gibi but the inlineing of _rebuild_cyborg_arq seems separate
15:20:24 sean-k-mooney i think i reverted that in my rebuild patch for a reason
15:21:31 gibi I think earlier _rebuild_cyborg_arq was extended to handle unshelve too but I asked not to do that
15:23:08 sean-k-mooney oh they are moveing wherre we get the nodename
15:23:51 gibi sean-k-mooney: in different action we have different source of the node
15:23:51 sean-k-mooney yah i suspec there is a second caller to _create_and_bind_arq_for_instance
15:24:04 sean-k-mooney that only has the hostname not the host
15:24:48 sean-k-mooney here https://review.opendev.org/c/openstack/nova/+/729563/26/nova/conductor/manager.py#1251
15:31:47 sean-k-mooney gibi: ya im not sure im a fan of folding _rebuild_cyborg_arq into the calling function
15:32:00 gibi that could be a comment
15:32:26 sean-k-mooney the other changes all look correct to me in the conductor manager however
15:33:50 sean-k-mooney well it only had one caller
15:34:21 sean-k-mooney and i intoduced it so ti could be reused for the other opertion so if we are not doing that removing it makes sense
15:34:41 sean-k-mooney its just the calling function is already quite large so its was nice to have it broken out
15:36:32 gibi yeah
15:36:41 gibi they tried to reuse it but that would neede a new flag
15:36:46 gibi so I was against that
15:36:54 gibi but we can keep it the hlper
15:36:56 gibi helper
15:44:12 sean-k-mooney its a nice to have rather then anything functional so not enough to -1 over
17:12:42 openstackgerrit Dan Smith proposed openstack/nova stable/victoria: Warn when starting services with older than N-1 computes https://review.opendev.org/c/openstack/nova/+/761923
17:13:19 dansmith gibi: if you will double-check my edits here I can just +2+W this as I just fixed nits ^
17:14:42 gibi dansmith: looking
17:16:24 gibi dansmith: thanks, looks good to me
17:20:38 dansmith ack
23:59:24 openstackgerrit Merged openstack/nova master: Improving the description for unshelve request body https://review.opendev.org/c/openstack/nova/+/767251
#openstack-nova - 2021-01-06
01:17:27 openstackgerrit melanie witt proposed openstack/nova master: Enable test_volume_backed_live_migration in tempest https://review.opendev.org/c/openstack/nova/+/528104
01:30:11 brinzhang0 bauzas, sean-k-mooney: I have replied the cyborg shelve/unshleve patch, pls review again, and I think there is no need to update in, one comments need to move the comments, and I will addressed in the followed up patch
06:53:47 openstackgerrit Mamduh proposed openstack/os-vif stable/queens: Refactor code of linux_net to more cleaner and increase performace https://review.opendev.org/c/openstack/os-vif/+/765941
06:53:48 openstackgerrit Mamduh proposed openstack/os-vif stable/queens: Fix - os-vif fails to get the correct UpLink Representor https://review.opendev.org/c/openstack/os-vif/+/765983
08:21:22 openstackgerrit Balazs Gibizer proposed openstack/nova master: Remove dead code from SchedulerReportClient https://review.opendev.org/c/openstack/nova/+/769467
09:42:14 alexe9191 Hello There,
09:42:50 alexe9191 I was wondering if I can run a newer version of nova-api/scheduler/conductor on an older version of nova-compute. Nova-compute would be kili for instance and the newer version of the API services would be on Pike or Rocky for instance?
09:44:47 gibi alexe9191: nova keeps compatibility between version N+1 controller and version N compute service
09:45:28 gibi so you can have Rocky controller services and Pike compute services
09:45:41 gibi sorry, so you can have Rocky controller services and Queens compute serivces
09:46:14 gibi as Q + 1 = R
09:46:18 alexe9191 interesting, what out of curiosity. Does this also work the other way around?
09:46:30 alexe9191 so nova-compute R and nova-api Q ?
09:49:19 gibi no, the controller services need to be upgraded first
09:50:17 alexe9191 I am also guessing that I can't immediately upgrade the nova-compute from kilo to rocky in one go? This also has to be done in a rolling fashion? So basically, K -> L -> etc > R on the nova-computes ?
09:51:38 gibi alexe9191: what we test is upstream is rolling upgrade as you guessed. There are some support for fast forward upgrade where you move more than one release forward in single step
09:51:52 gibi https://wiki.openstack.org/wiki/Fast_forward_upgrades
09:54:36 alexe9191 I honestly fail to understand that a little bit as the documentation is not really clear on fast forward upgrades
09:55:41 sean-k-mooney alexe9191: fast forward upgrades are not really a feature of the comonet project more an installer feature that the compoent project like nova try to not intentionally break
09:55:58 sean-k-mooney stictly speaking nova only support n to n+1
09:56:19 alexe9191 So basically this fast forward *feature* is more of a wrapper around what nova support ?
09:57:05 sean-k-mooney not quite it a way of runnign all the db updates and migration for the intermidat release so that you can skip them
09:57:42 sean-k-mooney on the nova side we try to keep that logic around for multiple release so that if you skip 3 version of the compute agent it will still have the compatibley code
09:58:09 sean-k-mooney so nova "supports" fast forward upgrades
09:58:18 alexe9191 So for instance running the DB migration scripts from newton while one is running Kilo ?
09:58:33 sean-k-mooney by keeping the compat code for several release beyond when it was striclty reuired
09:59:13 sean-k-mooney when you do an FFU there is a period where the contolplane is in accessable.
09:59:41 sean-k-mooney basiclly you move the contoler services to newton while the computes are on kilo
09:59:48 sean-k-mooney then you move the compute to newton
10:00:11 sean-k-mooney when you do that in principal the newton contoler cannot manage the kilo compute but the vms keep running
10:00:35 sean-k-mooney when you bring the compute back up to newton manageablity is restored
10:01:22 sean-k-mooney in pratice if there has been no RPC major version bump between the two version you have deploy the contoler could tehoretically mange compute more then 1 releasee old
10:01:27 sean-k-mooney but we dont test that
10:01:40 alexe9191 ok, and rolling upgrade being moving everything at once from one release to the other, both the APi & the nova-compute agents, K, L, N, ...
10:01:46 sean-k-mooney there is typically 5-6 release betwwen major rpc version bumps
10:02:32 alexe9191 That's very helpful to know!
10:02:50 sean-k-mooney alexe9191: rolling upgrades are where you move the contolplane in one go to n+1, then move the computes in a "rolling" or incremental fashion 1 by 1
10:03:31 sean-k-mooney rolling upgrades allow you to maintain managmeablity of the computes with mixed verions
10:03:41 sean-k-mooney but we only test a delta of 1 verions
10:04:11 sean-k-mooney we write the code such that it could be a larger delta but the test matix is too large for use to offialy support that
10:04:12 alexe9191 Right, so if upgrading K to L. I could basically roll the upgrade to the compute nodes one rack at a time from K to L
10:04:28 sean-k-mooney yes

Earlier   Later