| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-05-10 | |||
| 17:33:54 | sean-k-mooney | ok | |
| 17:34:45 | dansmith | the ossa went to the ML with more summary which might be good reading | |
| 17:35:01 | sean-k-mooney | ill read that after | |
| 17:35:12 | sean-k-mooney | we debated making the service user required in the past | |
| 17:35:26 | sean-k-mooney | not that it is we can use that for other usecases going forward | |
| 17:35:32 | dansmith | yup | |
| 17:35:42 | sean-k-mooney | like the manila stuff | |
| 17:35:57 | dansmith | yup, it was hard not to mention this during that discussion :) | |
| 17:36:11 | dansmith | seems like we could extend that to neutron and avoid nova having to be admin too | |
| 17:36:12 | sean-k-mooney | or any other interaction with other service where we need service to service interaction | |
| 17:36:18 | dansmith | yup | |
| 17:36:25 | sean-k-mooney | yep | |
| 17:36:43 | sean-k-mooney | we could simply nova config ot just need service user and no other sections for other services | |
| 17:36:50 | sean-k-mooney | at least by defualt | |
| 17:36:53 | dansmith | it's unfortunate to have to force that on a bunch of old deployments without a grace period, but.. not much choice | |
| 17:37:23 | sean-k-mooney | for what its worth i think many installer tools started using the service user a few years ago by default | |
| 17:37:37 | dansmith | yeah tripleo did | |
| 17:37:41 | sean-k-mooney | i think downstream its on by default in 17/train | |
| 17:37:50 | dansmith | and 16 | |
| 17:37:57 | dansmith | but it's not just the service user that is required, but also the role on that user | |
| 17:38:05 | sean-k-mooney | sorry 16/train ya | |
| 17:38:24 | sean-k-mooney | OSA and kolla have supprot but i dont know if its on by default | |
| 17:38:46 | dansmith | s'gonna be soon :) | |
| 17:39:54 | sean-k-mooney | unfortunetly it looks like no but looking at there config | |
| 17:40:11 | melwitt | oh huh, is stable/wallaby ci known to not work? | |
| 17:40:30 | sean-k-mooney | when the service user is not configred it may have been possibel to fallback to the info form the keystone authtoken section | |
| 17:40:52 | gmann | melwitt: what is failing, it should be green | |
| 17:40:54 | sean-k-mooney | melwitt: no but devstack might not have service users configred htere | |
| 17:41:12 | melwitt | gmann: ERROR: Could not find a version that satisfies the requirement tempest>=30.0.0 (from cinder-tempest-plugin) https://zuul.opendev.org/t/openstack/build/4c8a3d86710a4d50871613e2e578831a | |
| 17:41:54 | gmann | melwitt: oh, we did recent change there, let me check if cinder-tempest-pluing is not pinned there. we did pin the tempest. its incompatible tempest and c-t-p | |
| 17:42:25 | dansmith | incompatible because of me | |
| 17:43:25 | gmann | dansmith: no, you just exposed that but we should pin all plugin along with tempest to have such issue coming in future. | |
| 17:43:38 | dansmith | \o/ | |
| 17:44:31 | gmann | melwitt: dansmith: ah I pushed change while pinning tempest but did not realize it is not merged https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/871920 | |
| 17:44:59 | gmann | and that was the only change left for tempest pin on wallaby, :( https://review.opendev.org/q/topic:wallaby-pin-tempest | |
| 17:45:09 | melwitt | ah ok | |
| 17:46:50 | melwitt | I shan't recheck those then :) | |
| 17:48:47 | gmann | melwitt: yeah, it will fail until 871920 is merged, I will check it but not sure if ceph testing on wallaby is all working or not. | |
| 17:49:01 | melwitt | the plugin patch failed 25 tests ... not sure if they're unrelated fails | |
| 17:49:27 | gmann | melwitt: if urgent as CVE fix then we can make ceph job as n-v in wallaby for now to merge your fix | |
| 17:49:47 | gmann | yeah, I am not very positive of everything on ceph will work in wallaby | |
| 17:50:19 | melwitt | gmann: wallaby is nice-to-have I think, I just proposed it bc there was no change needed from the xena patch | |
| 17:51:03 | gmann | k | |
| 18:42:23 | dansmith | so I just noticed we're running a n-v job in the gate queue.. the ceph migration one | |
| 19:04:50 | dansmith | sean-k-mooney: so I've seen something that looks like this a number of times recently: https://92f377bbda976b6da17f-b08d635a9941444d91f5e0463ea7a01d.ssl.cf2.rackcdn.com/882852/1/gate/nova-grenade-multinode/b8a7af9/testr_results.html | |
| 19:05:40 | sean-k-mooney | test_list_migrations_in_flavor_resize_situation thats an interesting testname... | |
| 19:06:02 | dansmith | melwitt: dammit, ceph job timed out | |
| 19:06:03 | sean-k-mooney | hum binding failed | |
| 19:06:08 | dansmith | thought we were going to win the lottery | |
| 19:06:23 | melwitt | aw man | |
| 19:06:27 | sean-k-mooney | dansmith: the nv one or the working one. i guess the later | |
| 19:06:39 | melwitt | I've seen the port binding fail one a lot recently | |
| 19:06:44 | dansmith | the important one | |
| 19:07:14 | dansmith | we might want to up the timeout on the ceph job because it definitely takes legitimately longer with all the validations | |
| 19:07:32 | sean-k-mooney | so binding_failed is generally a rare thing so if its happen alto recently that concerning | |
| 19:07:39 | dansmith | and it's still steaming along when it is getting close to its timeout | |
| 19:07:51 | dansmith | I've seen it a number of times lately | |
| 19:08:01 | sean-k-mooney | whats it at currenlty | |
| 19:08:10 | sean-k-mooney | the timeout | |
| 19:08:21 | sean-k-mooney | 2 hours? | |
| 19:08:22 | dansmith | just over two hours | |
| 19:08:48 | sean-k-mooney | we can bump it to 150 mins i guess | |
| 19:09:13 | sean-k-mooney | i like it when test fail and give you the requst id | |
| 19:09:41 | sean-k-mooney | im going to see why that failed on the neutron side and ill let you know if i find anyting | |
| 19:12:13 | opendevreview | Dan Smith proposed openstack/nova master: Bump nova-ceph-multstore timeout https://review.opendev.org/c/openstack/nova/+/882890 | |
| 19:13:27 | sean-k-mooney | oh that is in seconds https://zuul-ci.org/docs/zuul/latest/job-content.html#var-zuul.timeout | |
| 23:02:53 | dansmith | here we go | |
| 23:05:36 | opendevreview | Merged openstack/nova master: Use force=True for os-brick disconnect during delete https://review.opendev.org/c/openstack/nova/+/882847 | |
| 23:05:43 | opendevreview | Merged openstack/nova master: Enable use of service user token with admin context https://review.opendev.org/c/openstack/nova/+/882852 | |
| #openstack-nova - 2023-05-11 | |||
| 00:19:29 | dansmith | melwitt: oh jeez I didn't realize that wasn't +V from zuul, thanks | |
| 00:20:28 | melwitt | np. it's hard to see through all the fail | |
| 00:20:54 | dansmith | just too many patches really | |
| 00:23:21 | melwitt | huh.. it looks like it's already been running. it was already +W by you.. I didn't think of the +W kicking off the job again /facepalm | |
| 00:23:53 | dansmith | oh I swore it was | |
| 00:24:02 | dansmith | I see your recheck put it back in the arm queue, | |
| 00:24:16 | dansmith | which is only a few minutes old but the other one has been running in regular check for over an hour | |
| 00:24:32 | melwitt | yeah. it doesn't restart the normal jobs thankfully | |
| 00:24:32 | dansmith | I thought I must have been looking at the wrong on | |
| 00:24:49 | melwitt | no, I'm just le dumb | |
| 00:25:10 | dansmith | no, there are too many identical patches :) | |
| 00:33:57 | dansmith | crap, 2023.1 patch is going to fail | |
| 00:34:23 | dansmith | dhcp fail in the guest | |
| 00:34:24 | dansmith | le sigh | |
| 00:34:28 | melwitt | dangit | |
| 00:49:17 | dansmith | top patch is +V waiting for another trip to gateland | |
| 00:49:34 | dansmith | man cripes | |
| 00:49:44 | dansmith | my friggin timeout patch is gonna fail again | |
| 00:50:24 | melwitt | 😑 | |
| 00:50:24 | melwitt | 😑 | |
| 00:50:39 | dansmith | [ 10.258086] Buffer I/O error on dev vda1, logical block 20, lost async page write | |
| 00:50:39 | dansmith | clearly something ceph related: | |
| 00:50:44 | dansmith | vda is the root disk, which is on ceph | |
| 00:51:05 | dansmith | last time around cinder just stopped creating volumes halfway through (maybe for similar reasons) | |
| 00:52:25 | dansmith | Out of memory: Killed process 55722 (ceph-osd) | |
| 00:52:28 | dansmith | that'll do it like every time | |
| 00:52:32 | dansmith | cripes | |
| 00:53:03 | melwitt | ouch | |
| 00:54:42 | dansmith | ah, | |
| 00:54:59 | dansmith | we're lowering the swap on our fatter job than the ceph job I increased in their repo | |
| 00:56:25 | opendevreview | Dan Smith proposed openstack/nova master: Bump nova-ceph-multstore timeout https://review.opendev.org/c/openstack/nova/+/882890 | |