| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-05-10 | |||
| 17:40:11 | melwitt | oh huh, is stable/wallaby ci known to not work? | |
| 17:40:30 | sean-k-mooney | when the service user is not configred it may have been possibel to fallback to the info form the keystone authtoken section | |
| 17:40:52 | gmann | melwitt: what is failing, it should be green | |
| 17:40:54 | sean-k-mooney | melwitt: no but devstack might not have service users configred htere | |
| 17:41:12 | melwitt | gmann: ERROR: Could not find a version that satisfies the requirement tempest>=30.0.0 (from cinder-tempest-plugin) https://zuul.opendev.org/t/openstack/build/4c8a3d86710a4d50871613e2e578831a | |
| 17:41:54 | gmann | melwitt: oh, we did recent change there, let me check if cinder-tempest-pluing is not pinned there. we did pin the tempest. its incompatible tempest and c-t-p | |
| 17:42:25 | dansmith | incompatible because of me | |
| 17:43:25 | gmann | dansmith: no, you just exposed that but we should pin all plugin along with tempest to have such issue coming in future. | |
| 17:43:38 | dansmith | \o/ | |
| 17:44:31 | gmann | melwitt: dansmith: ah I pushed change while pinning tempest but did not realize it is not merged https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/871920 | |
| 17:44:59 | gmann | and that was the only change left for tempest pin on wallaby, :( https://review.opendev.org/q/topic:wallaby-pin-tempest | |
| 17:45:09 | melwitt | ah ok | |
| 17:46:50 | melwitt | I shan't recheck those then :) | |
| 17:48:47 | gmann | melwitt: yeah, it will fail until 871920 is merged, I will check it but not sure if ceph testing on wallaby is all working or not. | |
| 17:49:01 | melwitt | the plugin patch failed 25 tests ... not sure if they're unrelated fails | |
| 17:49:27 | gmann | melwitt: if urgent as CVE fix then we can make ceph job as n-v in wallaby for now to merge your fix | |
| 17:49:47 | gmann | yeah, I am not very positive of everything on ceph will work in wallaby | |
| 17:50:19 | melwitt | gmann: wallaby is nice-to-have I think, I just proposed it bc there was no change needed from the xena patch | |
| 17:51:03 | gmann | k | |
| 18:42:23 | dansmith | so I just noticed we're running a n-v job in the gate queue.. the ceph migration one | |
| 19:04:50 | dansmith | sean-k-mooney: so I've seen something that looks like this a number of times recently: https://92f377bbda976b6da17f-b08d635a9941444d91f5e0463ea7a01d.ssl.cf2.rackcdn.com/882852/1/gate/nova-grenade-multinode/b8a7af9/testr_results.html | |
| 19:05:40 | sean-k-mooney | test_list_migrations_in_flavor_resize_situation thats an interesting testname... | |
| 19:06:02 | dansmith | melwitt: dammit, ceph job timed out | |
| 19:06:03 | sean-k-mooney | hum binding failed | |
| 19:06:08 | dansmith | thought we were going to win the lottery | |
| 19:06:23 | melwitt | aw man | |
| 19:06:27 | sean-k-mooney | dansmith: the nv one or the working one. i guess the later | |
| 19:06:39 | melwitt | I've seen the port binding fail one a lot recently | |
| 19:06:44 | dansmith | the important one | |
| 19:07:14 | dansmith | we might want to up the timeout on the ceph job because it definitely takes legitimately longer with all the validations | |
| 19:07:32 | sean-k-mooney | so binding_failed is generally a rare thing so if its happen alto recently that concerning | |
| 19:07:39 | dansmith | and it's still steaming along when it is getting close to its timeout | |
| 19:07:51 | dansmith | I've seen it a number of times lately | |
| 19:08:01 | sean-k-mooney | whats it at currenlty | |
| 19:08:10 | sean-k-mooney | the timeout | |
| 19:08:21 | sean-k-mooney | 2 hours? | |
| 19:08:22 | dansmith | just over two hours | |
| 19:08:48 | sean-k-mooney | we can bump it to 150 mins i guess | |
| 19:09:13 | sean-k-mooney | i like it when test fail and give you the requst id | |
| 19:09:41 | sean-k-mooney | im going to see why that failed on the neutron side and ill let you know if i find anyting | |
| 19:12:13 | opendevreview | Dan Smith proposed openstack/nova master: Bump nova-ceph-multstore timeout https://review.opendev.org/c/openstack/nova/+/882890 | |
| 19:13:27 | sean-k-mooney | oh that is in seconds https://zuul-ci.org/docs/zuul/latest/job-content.html#var-zuul.timeout | |
| 23:02:53 | dansmith | here we go | |
| 23:05:36 | opendevreview | Merged openstack/nova master: Use force=True for os-brick disconnect during delete https://review.opendev.org/c/openstack/nova/+/882847 | |
| 23:05:43 | opendevreview | Merged openstack/nova master: Enable use of service user token with admin context https://review.opendev.org/c/openstack/nova/+/882852 | |
| #openstack-nova - 2023-05-11 | |||
| 00:19:29 | dansmith | melwitt: oh jeez I didn't realize that wasn't +V from zuul, thanks | |
| 00:20:28 | melwitt | np. it's hard to see through all the fail | |
| 00:20:54 | dansmith | just too many patches really | |
| 00:23:21 | melwitt | huh.. it looks like it's already been running. it was already +W by you.. I didn't think of the +W kicking off the job again /facepalm | |
| 00:23:53 | dansmith | oh I swore it was | |
| 00:24:02 | dansmith | I see your recheck put it back in the arm queue, | |
| 00:24:16 | dansmith | which is only a few minutes old but the other one has been running in regular check for over an hour | |
| 00:24:32 | dansmith | I thought I must have been looking at the wrong on | |
| 00:24:32 | melwitt | yeah. it doesn't restart the normal jobs thankfully | |
| 00:24:49 | melwitt | no, I'm just le dumb | |
| 00:25:10 | dansmith | no, there are too many identical patches :) | |
| 00:33:57 | dansmith | crap, 2023.1 patch is going to fail | |
| 00:34:23 | dansmith | dhcp fail in the guest | |
| 00:34:24 | dansmith | le sigh | |
| 00:34:28 | melwitt | dangit | |
| 00:49:17 | dansmith | top patch is +V waiting for another trip to gateland | |
| 00:49:34 | dansmith | man cripes | |
| 00:49:44 | dansmith | my friggin timeout patch is gonna fail again | |
| 00:50:24 | melwitt | 😑 | |
| 00:50:24 | melwitt | 😑 | |
| 00:50:39 | dansmith | clearly something ceph related: | |
| 00:50:39 | dansmith | [ 10.258086] Buffer I/O error on dev vda1, logical block 20, lost async page write | |
| 00:50:44 | dansmith | vda is the root disk, which is on ceph | |
| 00:51:05 | dansmith | last time around cinder just stopped creating volumes halfway through (maybe for similar reasons) | |
| 00:52:25 | dansmith | Out of memory: Killed process 55722 (ceph-osd) | |
| 00:52:28 | dansmith | that'll do it like every time | |
| 00:52:32 | dansmith | cripes | |
| 00:53:03 | melwitt | ouch | |
| 00:54:42 | dansmith | ah, | |
| 00:54:59 | dansmith | we're lowering the swap on our fatter job than the ceph job I increased in their repo | |
| 00:56:25 | opendevreview | Dan Smith proposed openstack/nova master: Bump nova-ceph-multstore timeout https://review.opendev.org/c/openstack/nova/+/882890 | |
| 00:56:33 | dansmith | melwitt: ^ | |
| 00:59:06 | melwitt | I don't understand that sentence 😆 | |
| 00:59:06 | melwitt | I don't understand that sentence 😆 | |
| 00:59:52 | melwitt | ok nevermind | |
| 00:59:55 | dansmith | wait, in my commit message or | |
| 01:00:18 | melwitt | yeah you explained what I didn't understand in your commit message | |
| 01:00:25 | dansmith | I bumped the ceph job to 8G while jammifiying and cephadmifying it, but we inherit from that and set it down to 4G, which makes no sense because we run even more stuff than they do | |
| 01:00:30 | dansmith | ack | |
| 05:41:49 | opendevreview | Merged openstack/nova stable/2023.1: Use force=True for os-brick disconnect during delete https://review.opendev.org/c/openstack/nova/+/882858 | |
| 05:47:56 | opendevreview | Amit Uniyal proposed openstack/nova stable/2023.1: Have host look for CPU controller of cgroupsv2 location. https://review.opendev.org/c/openstack/nova/+/882913 | |
| 05:50:06 | opendevreview | Amit Uniyal proposed openstack/nova stable/zed: Have host look for CPU controller of cgroupsv2 location. https://review.opendev.org/c/openstack/nova/+/882914 | |
| 06:43:51 | gibi | rechecked https://review.opendev.org/c/openstack/nova/+/882890 as it failed nova-next in the gate with unrealted timeouts | |
| 06:46:37 | gibi | also rechecked https://review.opendev.org/c/openstack/nova/+/882859 | |
| 06:47:29 | gibi | both failed with http read timeout on various openstack APIs (cinder, neutron, nova) | |
| 06:48:05 | gibi | I think this is tracked here as a bug https://bugs.launchpad.net/tempest/+bug/1999893 | |
| 07:05:24 | bauzas | morning | |
| 07:05:45 | bauzas | gibi: catching up the world explosion after my yesterday PTO | |
| 07:29:12 | gibi | I don't have the full context as the CVE got public why I was away | |
| 07:29:20 | gibi | I see the fixes proposed so I'm trying to land them | |
| 07:29:25 | opendevreview | Amit Uniyal proposed openstack/nova stable/yoga: Have host look for CPU controller of cgroupsv2 location. https://review.opendev.org/c/openstack/nova/+/882920 | |
| 07:32:10 | gibi | s/why/while/ | |
| 07:32:16 | opendevreview | Sylvain Bauza proposed openstack/nova stable/2023.1: Revert "Debug Nova APIs call failures" https://review.opendev.org/c/openstack/nova/+/882783 | |
| 07:33:07 | bauzas | gibi: np, just looking at gerrit reviews | |
| 07:33:14 | bauzas | gibi: do you need some explanations ? | |