Earlier  
Posted Nick Remark
#openstack-nova - 2023-03-03
10:20:12 bauzas https://github.com/openstack/nova/blob/master/nova/objects/service.py#L212-L215 and that's what was changed by the RPC version for the 60 service version
10:20:57 bauzas unless you deploy with master, of course
10:25:18 admin1 bauzas, so the servers that are 60 were not online ..
10:25:33 admin1 they were off temporarily
10:25:40 admin1 but this blocked the whole upgrade process
10:26:17 bauzas the compute state isn't and shouldn't be checked for safety reasons
10:26:56 admin1 it did .. temporarily what i did was update nova.services set version=61 where version=60 and trying to run the playbook again
10:27:12 admin1 if it works, then i am all good .. else i have to report it here again
10:27:27 admin1 if this works, then i can open a bug report saying unavailable compute node blocked upgrade
10:27:50 bauzas https://docs.openstack.org/nova/latest/cli/nova-status.html#nova-status-checks helps to test your upgrade
10:28:27 bauzas admin1: again, that's by design that we don't allow non-upgraded compute to be left registered
10:28:44 bauzas admin1: and that's why we have the workaround option for that intent
10:29:54 bauzas https://github.com/openstack/nova/blob/master/nova/cmd/status.py#L251
10:30:39 bauzas which exactly tests the support contract *before* you upgrade https://github.com/openstack/nova/blob/59f7a524fd4ded3c17b10abcedb0baff769c3a8a/nova/utils.py#L1052
11:07:51 sean-k-mooney admin1: havign the server be off is expected to block the upgrade process
11:08:04 sean-k-mooney that would not be a bug
11:08:40 sean-k-mooney because if the service verion is 60 it means they never started with the offical yoga release (61)
11:09:08 opendevreview Amit Uniyal proposed openstack/nova master: Allow swap resize from non-zero to zero https://review.opendev.org/c/openstack/nova/+/857339
11:09:11 sean-k-mooney if you had started them with yoga and then they were stop it would not cause this issue
11:09:43 admin1 i think they were off 2 days before when i did wallaby -> yoga
11:09:54 admin1 but wallaby -> yoga did not complained of this .. was this check added in zed ?
11:10:36 sean-k-mooney this was alwasy a requirement and we decied to start enforcining it in yoga becasue of operators violating the upgrade contract
11:10:41 sean-k-mooney and filing bugs :)
11:10:44 admin1 :D
11:11:37 sean-k-mooney nova before 2023.1/2024.1 only allows n to n+1 upgrades
11:11:54 sean-k-mooney we put the workaround option in place ans an escape hatch
11:12:42 sean-k-mooney so if you want to run nova in an unsupproted state you can but it should never be requried if you are upgrading withing the upgrade contract
11:13:02 admin1 i understand now .. will make sure no computes are down next time we upgrade
11:15:08 sean-k-mooney provided we do not do an rpc bump its generally possibel for > n->n+1 to function but the first time we tested that was yoga to antelope(2023.1) as a dry run for 2023.1->2024.1
11:16:57 sean-k-mooney we will be offically testing that going forward in case you are not aware of this chagne https://governance.openstack.org/tc/resolutions/20220210-release-cadence-adjustment.html
11:19:32 opendevreview Amit Uniyal proposed openstack/nova master: Allow swap resize from non-zero to zero https://review.opendev.org/c/openstack/nova/+/857339
14:24:11 dansmith bauzas: \o/
14:33:58 dansmith bauzas: are you cooking up the rc1 patch?
14:34:15 bauzas I was waiting for the revert to arrive
14:34:28 dansmith ah okay, I see it's close
14:34:40 dansmith sorry I didn't recheck that because I thought we were punting
14:34:44 bauzas shit no
14:34:46 bauzas https://zuul.openstack.org/status#873584
14:34:51 dansmith oh yep
14:34:59 bauzas post_failure
14:35:04 bauzas ok, so I'll skip it
14:35:20 bauzas gibi: sean-k-mooney: for your sake of knowledge, I'm gonna branch RC1 without the logging revert
14:35:39 bauzas we'll backport the revert later aftr GA
14:36:51 bauzas hmmmm
14:37:17 bauzas dansmith: actually, it looks like the release team agrees us some graceful extra period for branching RC1
14:37:44 bauzas (they're on meeting now)
14:37:56 dansmith okay
14:38:03 gibi the revert is in the gate queue now
14:38:11 sean-k-mooney ok we modified it ot be safe in production even so having it in RC1 is not terible
14:38:12 bauzas gibi: yup, but failing
14:38:13 gibi so if we are lucky it might merge today
14:38:15 gibi ohh
14:38:19 gibi sh*t
14:38:29 bauzas we were so close
14:38:38 sean-k-mooney but we can ask them to reque it
14:38:43 bauzas I'll claim for a RC1 patch on Monday
14:38:54 sean-k-mooney if there is a long delay
14:38:57 bauzas and I'll recheck this revert by the next 4 mins
14:39:15 dansmith 18 other things in the gate right now
14:39:30 gibi bauzas: I can shepherd the patch during Saturday and a bit on Sunday as well.
14:39:35 dansmith so it'll be a bit if it re-runs, but it's also not a critical patch
14:40:03 sean-k-mooney post_failure form nova next. unfortunet
14:40:04 bauzas dansmith: I don't disagree
14:40:21 bauzas but it will be a bit of a pain to backport the revert if we go
14:40:49 bauzas if the release team says they're OK with releasing on Monday, then meh, we gonna try this weekend
14:41:00 bauzas gibi: last time you were way luckier than me
14:42:48 sean-k-mooney we technially didnt run out of memory but it got pretty clsoe memory_tracker low_point: 730
14:43:01 sean-k-mooney * memory_tracker low_point: 7308
14:43:16 dansmith sean-k-mooney: has nova-next been OOMing?
14:43:26 sean-k-mooney MemAvailable: 9152 kB
14:43:51 sean-k-mooney i think its been surviing because of swap
14:44:02 sean-k-mooney Mar 03 13:53:10.114492 np0033355853 memory_tracker.sh[131948]: SwapTotal: 4194300 kB
14:44:04 sean-k-mooney Mar 03 13:53:10.114492 np0033355853 memory_tracker.sh[131948]: SwapFree: 0 kB
14:44:18 sean-k-mooney https://zuul.opendev.org/t/openstack/build/ec34b5fa7a354e19a6919d167268cb8b/log/controller/logs/screen-memory_tracker.txt#3011
14:44:59 sean-k-mooney dansmith: are you wondering if this is related to the mariadb tweaks ye did
14:45:31 sean-k-mooney keystone was giving 503s
14:45:33 dansmith sean-k-mooney: those tweaks are disabled by default in devstack right now
14:45:51 dansmith I'm just saying if we're memory constrained on that job, we might want to enable those tweaks
14:45:57 sean-k-mooney and when i see that it often because of the db/service getting oom killed
14:46:02 dansmith it seem to have done well for the ceph one
14:46:03 sean-k-mooney yep
14:46:23 sean-k-mooney im just looking to see if i can confim that in the logs
14:46:45 sean-k-mooney but that is why i was checkign the memory tracker i think we are runnign very close to out of memory if we have not hit it
14:47:11 dansmith on the ceph job my tweaks dropped mysql to half of what it was using (~800m to ~400m)
14:47:23 sean-k-mooney ar 03 13:53:09 np0033355853 kernel: sshd invoked oom-killer: gfp_mask=0x1100cca(GFP_HIGHUSER_MOVABLE), order=0, oom_score_adj=0
14:47:34 bauzas https://zuul.opendev.org/t/openstack/build/ec34b5fa7a354e19a6919d167268cb8b
14:47:34 dansmith but, less memory usage could impair performance and make other things worse of course
14:47:48 sean-k-mooney so yes we are
14:47:50 sean-k-mooney https://zuul.opendev.org/t/openstack/build/ec34b5fa7a354e19a6919d167268cb8b/log/controller/logs/syslog.txt#5871
14:47:56 bauzas we had two problems
14:48:01 bauzas a unresponsive API
14:48:11 bauzas and some leaked allocs
14:48:42 sean-k-mooney the api issue are because mysql got killed
14:48:45 sean-k-mooney Mar 03 13:53:09 np0033355853 kernel: Out of memory: Killed process 47910 (mysqld) total-vm:5223564kB, anon-rss:328112kB, file-rss:0kB, shmem-rss:0kB, UID:116 pgtables:2648kB oom_score_adj:0
14:49:04 dansmith yeah that's usually how it works
14:49:16 dansmith mysql is killed and then we stop being able to talk to keystone (et al)
14:49:36 sean-k-mooney yep
14:50:10 sean-k-mooney so 1 we shoudl enabel that devstack feature for nova-next 2 we shoudl consider doing it by default
14:50:26 sean-k-mooney assuming it does not regress over all job time too much
14:50:40 dansmith we were going to do it by default after everyone is branched, to see if it is reasonable across the board, but not to break anyone before release

Earlier   Later