Earlier  
Posted Nick Remark
#openstack-nova - 2023-03-03
10:14:30 admin1 output of select host,version from nova.services => https://gist.githubusercontent.com/a1git/13ceb2e181dab9532a5b229b2915b478/raw/eee560226905d21159808f36a038c9201d8e2d23/gistfile1.txt
10:14:37 admin1 i think some are affected
10:18:53 bauzas admin1: what's strange is that 60 is an interim service version
10:19:11 bauzas did you use a milestone or something for some computes ?
10:19:43 bauzas https://github.com/openstack/nova/blob/master/nova/objects/service.py#L260-L261 Xena is 57 and Yoga is 61
10:20:12 bauzas https://github.com/openstack/nova/blob/master/nova/objects/service.py#L212-L215 and that's what was changed by the RPC version for the 60 service version
10:20:57 bauzas unless you deploy with master, of course
10:25:18 admin1 bauzas, so the servers that are 60 were not online ..
10:25:33 admin1 they were off temporarily
10:25:40 admin1 but this blocked the whole upgrade process
10:26:17 bauzas the compute state isn't and shouldn't be checked for safety reasons
10:26:56 admin1 it did .. temporarily what i did was update nova.services set version=61 where version=60 and trying to run the playbook again
10:27:12 admin1 if it works, then i am all good .. else i have to report it here again
10:27:27 admin1 if this works, then i can open a bug report saying unavailable compute node blocked upgrade
10:27:50 bauzas https://docs.openstack.org/nova/latest/cli/nova-status.html#nova-status-checks helps to test your upgrade
10:28:27 bauzas admin1: again, that's by design that we don't allow non-upgraded compute to be left registered
10:28:44 bauzas admin1: and that's why we have the workaround option for that intent
10:29:54 bauzas https://github.com/openstack/nova/blob/master/nova/cmd/status.py#L251
10:30:39 bauzas which exactly tests the support contract *before* you upgrade https://github.com/openstack/nova/blob/59f7a524fd4ded3c17b10abcedb0baff769c3a8a/nova/utils.py#L1052
11:07:51 sean-k-mooney admin1: havign the server be off is expected to block the upgrade process
11:08:04 sean-k-mooney that would not be a bug
11:08:40 sean-k-mooney because if the service verion is 60 it means they never started with the offical yoga release (61)
11:09:08 opendevreview Amit Uniyal proposed openstack/nova master: Allow swap resize from non-zero to zero https://review.opendev.org/c/openstack/nova/+/857339
11:09:11 sean-k-mooney if you had started them with yoga and then they were stop it would not cause this issue
11:09:43 admin1 i think they were off 2 days before when i did wallaby -> yoga
11:09:54 admin1 but wallaby -> yoga did not complained of this .. was this check added in zed ?
11:10:36 sean-k-mooney this was alwasy a requirement and we decied to start enforcining it in yoga becasue of operators violating the upgrade contract
11:10:41 sean-k-mooney and filing bugs :)
11:10:44 admin1 :D
11:11:37 sean-k-mooney nova before 2023.1/2024.1 only allows n to n+1 upgrades
11:11:54 sean-k-mooney we put the workaround option in place ans an escape hatch
11:12:42 sean-k-mooney so if you want to run nova in an unsupproted state you can but it should never be requried if you are upgrading withing the upgrade contract
11:13:02 admin1 i understand now .. will make sure no computes are down next time we upgrade
11:15:08 sean-k-mooney provided we do not do an rpc bump its generally possibel for > n->n+1 to function but the first time we tested that was yoga to antelope(2023.1) as a dry run for 2023.1->2024.1
11:16:57 sean-k-mooney we will be offically testing that going forward in case you are not aware of this chagne https://governance.openstack.org/tc/resolutions/20220210-release-cadence-adjustment.html
11:19:32 opendevreview Amit Uniyal proposed openstack/nova master: Allow swap resize from non-zero to zero https://review.opendev.org/c/openstack/nova/+/857339
14:24:11 dansmith bauzas: \o/
14:33:58 dansmith bauzas: are you cooking up the rc1 patch?
14:34:15 bauzas I was waiting for the revert to arrive
14:34:28 dansmith ah okay, I see it's close
14:34:40 dansmith sorry I didn't recheck that because I thought we were punting
14:34:44 bauzas shit no
14:34:46 bauzas https://zuul.openstack.org/status#873584
14:34:51 dansmith oh yep
14:34:59 bauzas post_failure
14:35:04 bauzas ok, so I'll skip it
14:35:20 bauzas gibi: sean-k-mooney: for your sake of knowledge, I'm gonna branch RC1 without the logging revert
14:35:39 bauzas we'll backport the revert later aftr GA
14:36:51 bauzas hmmmm
14:37:17 bauzas dansmith: actually, it looks like the release team agrees us some graceful extra period for branching RC1
14:37:44 bauzas (they're on meeting now)
14:37:56 dansmith okay
14:38:03 gibi the revert is in the gate queue now
14:38:11 sean-k-mooney ok we modified it ot be safe in production even so having it in RC1 is not terible
14:38:12 bauzas gibi: yup, but failing
14:38:13 gibi so if we are lucky it might merge today
14:38:15 gibi ohh
14:38:19 gibi sh*t
14:38:29 bauzas we were so close
14:38:38 sean-k-mooney but we can ask them to reque it
14:38:43 bauzas I'll claim for a RC1 patch on Monday
14:38:54 sean-k-mooney if there is a long delay
14:38:57 bauzas and I'll recheck this revert by the next 4 mins
14:39:15 dansmith 18 other things in the gate right now
14:39:30 gibi bauzas: I can shepherd the patch during Saturday and a bit on Sunday as well.
14:39:35 dansmith so it'll be a bit if it re-runs, but it's also not a critical patch
14:40:03 sean-k-mooney post_failure form nova next. unfortunet
14:40:04 bauzas dansmith: I don't disagree
14:40:21 bauzas but it will be a bit of a pain to backport the revert if we go
14:40:49 bauzas if the release team says they're OK with releasing on Monday, then meh, we gonna try this weekend
14:41:00 bauzas gibi: last time you were way luckier than me
14:42:48 sean-k-mooney we technially didnt run out of memory but it got pretty clsoe memory_tracker low_point: 730
14:43:01 sean-k-mooney * memory_tracker low_point: 7308
14:43:16 dansmith sean-k-mooney: has nova-next been OOMing?
14:43:26 sean-k-mooney MemAvailable: 9152 kB
14:43:51 sean-k-mooney i think its been surviing because of swap
14:44:02 sean-k-mooney Mar 03 13:53:10.114492 np0033355853 memory_tracker.sh[131948]: SwapTotal: 4194300 kB
14:44:04 sean-k-mooney Mar 03 13:53:10.114492 np0033355853 memory_tracker.sh[131948]: SwapFree: 0 kB
14:44:18 sean-k-mooney https://zuul.opendev.org/t/openstack/build/ec34b5fa7a354e19a6919d167268cb8b/log/controller/logs/screen-memory_tracker.txt#3011
14:44:59 sean-k-mooney dansmith: are you wondering if this is related to the mariadb tweaks ye did
14:45:31 sean-k-mooney keystone was giving 503s
14:45:33 dansmith sean-k-mooney: those tweaks are disabled by default in devstack right now
14:45:51 dansmith I'm just saying if we're memory constrained on that job, we might want to enable those tweaks
14:45:57 sean-k-mooney and when i see that it often because of the db/service getting oom killed
14:46:02 dansmith it seem to have done well for the ceph one
14:46:03 sean-k-mooney yep
14:46:23 sean-k-mooney im just looking to see if i can confim that in the logs
14:46:45 sean-k-mooney but that is why i was checkign the memory tracker i think we are runnign very close to out of memory if we have not hit it
14:47:11 dansmith on the ceph job my tweaks dropped mysql to half of what it was using (~800m to ~400m)
14:47:23 sean-k-mooney ar 03 13:53:09 np0033355853 kernel: sshd invoked oom-killer: gfp_mask=0x1100cca(GFP_HIGHUSER_MOVABLE), order=0, oom_score_adj=0
14:47:34 dansmith but, less memory usage could impair performance and make other things worse of course
14:47:34 bauzas https://zuul.opendev.org/t/openstack/build/ec34b5fa7a354e19a6919d167268cb8b
14:47:48 sean-k-mooney so yes we are
14:47:50 sean-k-mooney https://zuul.opendev.org/t/openstack/build/ec34b5fa7a354e19a6919d167268cb8b/log/controller/logs/syslog.txt#5871
14:47:56 bauzas we had two problems
14:48:01 bauzas a unresponsive API
14:48:11 bauzas and some leaked allocs
14:48:42 sean-k-mooney the api issue are because mysql got killed
14:48:45 sean-k-mooney Mar 03 13:53:09 np0033355853 kernel: Out of memory: Killed process 47910 (mysqld) total-vm:5223564kB, anon-rss:328112kB, file-rss:0kB, shmem-rss:0kB, UID:116 pgtables:2648kB oom_score_adj:0
14:49:04 dansmith yeah that's usually how it works

Earlier   Later