| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-03-03 | |||
| 10:20:12 | bauzas | https://github.com/openstack/nova/blob/master/nova/objects/service.py#L212-L215 and that's what was changed by the RPC version for the 60 service version | |
| 10:20:57 | bauzas | unless you deploy with master, of course | |
| 10:25:18 | admin1 | bauzas, so the servers that are 60 were not online .. | |
| 10:25:33 | admin1 | they were off temporarily | |
| 10:25:40 | admin1 | but this blocked the whole upgrade process | |
| 10:26:17 | bauzas | the compute state isn't and shouldn't be checked for safety reasons | |
| 10:26:56 | admin1 | it did .. temporarily what i did was update nova.services set version=61 where version=60 and trying to run the playbook again | |
| 10:27:12 | admin1 | if it works, then i am all good .. else i have to report it here again | |
| 10:27:27 | admin1 | if this works, then i can open a bug report saying unavailable compute node blocked upgrade | |
| 10:27:50 | bauzas | https://docs.openstack.org/nova/latest/cli/nova-status.html#nova-status-checks helps to test your upgrade | |
| 10:28:27 | bauzas | admin1: again, that's by design that we don't allow non-upgraded compute to be left registered | |
| 10:28:44 | bauzas | admin1: and that's why we have the workaround option for that intent | |
| 10:29:54 | bauzas | https://github.com/openstack/nova/blob/master/nova/cmd/status.py#L251 | |
| 10:30:39 | bauzas | which exactly tests the support contract *before* you upgrade https://github.com/openstack/nova/blob/59f7a524fd4ded3c17b10abcedb0baff769c3a8a/nova/utils.py#L1052 | |
| 11:07:51 | sean-k-mooney | admin1: havign the server be off is expected to block the upgrade process | |
| 11:08:04 | sean-k-mooney | that would not be a bug | |
| 11:08:40 | sean-k-mooney | because if the service verion is 60 it means they never started with the offical yoga release (61) | |
| 11:09:08 | opendevreview | Amit Uniyal proposed openstack/nova master: Allow swap resize from non-zero to zero https://review.opendev.org/c/openstack/nova/+/857339 | |
| 11:09:11 | sean-k-mooney | if you had started them with yoga and then they were stop it would not cause this issue | |
| 11:09:43 | admin1 | i think they were off 2 days before when i did wallaby -> yoga | |
| 11:09:54 | admin1 | but wallaby -> yoga did not complained of this .. was this check added in zed ? | |
| 11:10:36 | sean-k-mooney | this was alwasy a requirement and we decied to start enforcining it in yoga becasue of operators violating the upgrade contract | |
| 11:10:41 | sean-k-mooney | and filing bugs :) | |
| 11:10:44 | admin1 | :D | |
| 11:11:37 | sean-k-mooney | nova before 2023.1/2024.1 only allows n to n+1 upgrades | |
| 11:11:54 | sean-k-mooney | we put the workaround option in place ans an escape hatch | |
| 11:12:42 | sean-k-mooney | so if you want to run nova in an unsupproted state you can but it should never be requried if you are upgrading withing the upgrade contract | |
| 11:13:02 | admin1 | i understand now .. will make sure no computes are down next time we upgrade | |
| 11:15:08 | sean-k-mooney | provided we do not do an rpc bump its generally possibel for > n->n+1 to function but the first time we tested that was yoga to antelope(2023.1) as a dry run for 2023.1->2024.1 | |
| 11:16:57 | sean-k-mooney | we will be offically testing that going forward in case you are not aware of this chagne https://governance.openstack.org/tc/resolutions/20220210-release-cadence-adjustment.html | |
| 11:19:32 | opendevreview | Amit Uniyal proposed openstack/nova master: Allow swap resize from non-zero to zero https://review.opendev.org/c/openstack/nova/+/857339 | |
| 14:24:11 | dansmith | bauzas: \o/ | |
| 14:33:58 | dansmith | bauzas: are you cooking up the rc1 patch? | |
| 14:34:15 | bauzas | I was waiting for the revert to arrive | |
| 14:34:28 | dansmith | ah okay, I see it's close | |
| 14:34:40 | dansmith | sorry I didn't recheck that because I thought we were punting | |
| 14:34:44 | bauzas | shit no | |
| 14:34:46 | bauzas | https://zuul.openstack.org/status#873584 | |
| 14:34:51 | dansmith | oh yep | |
| 14:34:59 | bauzas | post_failure | |
| 14:35:04 | bauzas | ok, so I'll skip it | |
| 14:35:20 | bauzas | gibi: sean-k-mooney: for your sake of knowledge, I'm gonna branch RC1 without the logging revert | |
| 14:35:39 | bauzas | we'll backport the revert later aftr GA | |
| 14:36:51 | bauzas | hmmmm | |
| 14:37:17 | bauzas | dansmith: actually, it looks like the release team agrees us some graceful extra period for branching RC1 | |
| 14:37:44 | bauzas | (they're on meeting now) | |
| 14:37:56 | dansmith | okay | |
| 14:38:03 | gibi | the revert is in the gate queue now | |
| 14:38:11 | sean-k-mooney | ok we modified it ot be safe in production even so having it in RC1 is not terible | |
| 14:38:12 | bauzas | gibi: yup, but failing | |
| 14:38:13 | gibi | so if we are lucky it might merge today | |
| 14:38:15 | gibi | ohh | |
| 14:38:19 | gibi | sh*t | |
| 14:38:29 | bauzas | we were so close | |
| 14:38:38 | sean-k-mooney | but we can ask them to reque it | |
| 14:38:43 | bauzas | I'll claim for a RC1 patch on Monday | |
| 14:38:54 | sean-k-mooney | if there is a long delay | |
| 14:38:57 | bauzas | and I'll recheck this revert by the next 4 mins | |
| 14:39:15 | dansmith | 18 other things in the gate right now | |
| 14:39:30 | gibi | bauzas: I can shepherd the patch during Saturday and a bit on Sunday as well. | |
| 14:39:35 | dansmith | so it'll be a bit if it re-runs, but it's also not a critical patch | |
| 14:40:03 | sean-k-mooney | post_failure form nova next. unfortunet | |
| 14:40:04 | bauzas | dansmith: I don't disagree | |
| 14:40:21 | bauzas | but it will be a bit of a pain to backport the revert if we go | |
| 14:40:49 | bauzas | if the release team says they're OK with releasing on Monday, then meh, we gonna try this weekend | |
| 14:41:00 | bauzas | gibi: last time you were way luckier than me | |
| 14:42:48 | sean-k-mooney | we technially didnt run out of memory but it got pretty clsoe memory_tracker low_point: 730 | |
| 14:43:01 | sean-k-mooney | * memory_tracker low_point: 7308 | |
| 14:43:16 | dansmith | sean-k-mooney: has nova-next been OOMing? | |
| 14:43:26 | sean-k-mooney | MemAvailable: 9152 kB | |
| 14:43:51 | sean-k-mooney | i think its been surviing because of swap | |
| 14:44:02 | sean-k-mooney | Mar 03 13:53:10.114492 np0033355853 memory_tracker.sh[131948]: SwapTotal: 4194300 kB | |
| 14:44:04 | sean-k-mooney | Mar 03 13:53:10.114492 np0033355853 memory_tracker.sh[131948]: SwapFree: 0 kB | |
| 14:44:18 | sean-k-mooney | https://zuul.opendev.org/t/openstack/build/ec34b5fa7a354e19a6919d167268cb8b/log/controller/logs/screen-memory_tracker.txt#3011 | |
| 14:44:59 | sean-k-mooney | dansmith: are you wondering if this is related to the mariadb tweaks ye did | |
| 14:45:31 | sean-k-mooney | keystone was giving 503s | |
| 14:45:33 | dansmith | sean-k-mooney: those tweaks are disabled by default in devstack right now | |
| 14:45:51 | dansmith | I'm just saying if we're memory constrained on that job, we might want to enable those tweaks | |
| 14:45:57 | sean-k-mooney | and when i see that it often because of the db/service getting oom killed | |
| 14:46:02 | dansmith | it seem to have done well for the ceph one | |
| 14:46:03 | sean-k-mooney | yep | |
| 14:46:23 | sean-k-mooney | im just looking to see if i can confim that in the logs | |
| 14:46:45 | sean-k-mooney | but that is why i was checkign the memory tracker i think we are runnign very close to out of memory if we have not hit it | |
| 14:47:11 | dansmith | on the ceph job my tweaks dropped mysql to half of what it was using (~800m to ~400m) | |
| 14:47:23 | sean-k-mooney | ar 03 13:53:09 np0033355853 kernel: sshd invoked oom-killer: gfp_mask=0x1100cca(GFP_HIGHUSER_MOVABLE), order=0, oom_score_adj=0 | |
| 14:47:34 | dansmith | but, less memory usage could impair performance and make other things worse of course | |
| 14:47:34 | bauzas | https://zuul.opendev.org/t/openstack/build/ec34b5fa7a354e19a6919d167268cb8b | |
| 14:47:48 | sean-k-mooney | so yes we are | |
| 14:47:50 | sean-k-mooney | https://zuul.opendev.org/t/openstack/build/ec34b5fa7a354e19a6919d167268cb8b/log/controller/logs/syslog.txt#5871 | |
| 14:47:56 | bauzas | we had two problems | |
| 14:48:01 | bauzas | a unresponsive API | |
| 14:48:11 | bauzas | and some leaked allocs | |
| 14:48:42 | sean-k-mooney | the api issue are because mysql got killed | |
| 14:48:45 | sean-k-mooney | Mar 03 13:53:09 np0033355853 kernel: Out of memory: Killed process 47910 (mysqld) total-vm:5223564kB, anon-rss:328112kB, file-rss:0kB, shmem-rss:0kB, UID:116 pgtables:2648kB oom_score_adj:0 | |
| 14:49:04 | dansmith | yeah that's usually how it works | |
| 14:49:16 | dansmith | mysql is killed and then we stop being able to talk to keystone (et al) | |
| 14:49:36 | sean-k-mooney | yep | |
| 14:50:10 | sean-k-mooney | so 1 we shoudl enabel that devstack feature for nova-next 2 we shoudl consider doing it by default | |
| 14:50:26 | sean-k-mooney | assuming it does not regress over all job time too much | |
| 14:50:40 | dansmith | we were going to do it by default after everyone is branched, to see if it is reasonable across the board, but not to break anyone before release | |