Earlier  
Posted Nick Remark
#openstack-nova - 2023-05-10
17:36:12 sean-k-mooney or any other interaction with other service where we need service to service interaction
17:36:18 dansmith yup
17:36:25 sean-k-mooney yep
17:36:43 sean-k-mooney we could simply nova config ot just need service user and no other sections for other services
17:36:50 sean-k-mooney at least by defualt
17:36:53 dansmith it's unfortunate to have to force that on a bunch of old deployments without a grace period, but.. not much choice
17:37:23 sean-k-mooney for what its worth i think many installer tools started using the service user a few years ago by default
17:37:37 dansmith yeah tripleo did
17:37:41 sean-k-mooney i think downstream its on by default in 17/train
17:37:50 dansmith and 16
17:37:57 dansmith but it's not just the service user that is required, but also the role on that user
17:38:05 sean-k-mooney sorry 16/train ya
17:38:24 sean-k-mooney OSA and kolla have supprot but i dont know if its on by default
17:38:46 dansmith s'gonna be soon :)
17:39:54 sean-k-mooney unfortunetly it looks like no but looking at there config
17:40:11 melwitt oh huh, is stable/wallaby ci known to not work?
17:40:30 sean-k-mooney when the service user is not configred it may have been possibel to fallback to the info form the keystone authtoken section
17:40:52 gmann melwitt: what is failing, it should be green
17:40:54 sean-k-mooney melwitt: no but devstack might not have service users configred htere
17:41:12 melwitt gmann: ERROR: Could not find a version that satisfies the requirement tempest>=30.0.0 (from cinder-tempest-plugin) https://zuul.opendev.org/t/openstack/build/4c8a3d86710a4d50871613e2e578831a
17:41:54 gmann melwitt: oh, we did recent change there, let me check if cinder-tempest-pluing is not pinned there. we did pin the tempest. its incompatible tempest and c-t-p
17:42:25 dansmith incompatible because of me
17:43:25 gmann dansmith: no, you just exposed that but we should pin all plugin along with tempest to have such issue coming in future.
17:43:38 dansmith \o/
17:44:31 gmann melwitt: dansmith: ah I pushed change while pinning tempest but did not realize it is not merged https://review.opendev.org/c/openstack/devstack-plugin-ceph/+/871920
17:44:59 gmann and that was the only change left for tempest pin on wallaby, :( https://review.opendev.org/q/topic:wallaby-pin-tempest
17:45:09 melwitt ah ok
17:46:50 melwitt I shan't recheck those then :)
17:48:47 gmann melwitt: yeah, it will fail until 871920 is merged, I will check it but not sure if ceph testing on wallaby is all working or not.
17:49:01 melwitt the plugin patch failed 25 tests ... not sure if they're unrelated fails
17:49:27 gmann melwitt: if urgent as CVE fix then we can make ceph job as n-v in wallaby for now to merge your fix
17:49:47 gmann yeah, I am not very positive of everything on ceph will work in wallaby
17:50:19 melwitt gmann: wallaby is nice-to-have I think, I just proposed it bc there was no change needed from the xena patch
17:51:03 gmann k
18:42:23 dansmith so I just noticed we're running a n-v job in the gate queue.. the ceph migration one
19:04:50 dansmith sean-k-mooney: so I've seen something that looks like this a number of times recently: https://92f377bbda976b6da17f-b08d635a9941444d91f5e0463ea7a01d.ssl.cf2.rackcdn.com/882852/1/gate/nova-grenade-multinode/b8a7af9/testr_results.html
19:05:40 sean-k-mooney test_list_migrations_in_flavor_resize_situation thats an interesting testname...
19:06:02 dansmith melwitt: dammit, ceph job timed out
19:06:03 sean-k-mooney hum binding failed
19:06:08 dansmith thought we were going to win the lottery
19:06:23 melwitt aw man
19:06:27 sean-k-mooney dansmith: the nv one or the working one. i guess the later
19:06:39 melwitt I've seen the port binding fail one a lot recently
19:06:44 dansmith the important one
19:07:14 dansmith we might want to up the timeout on the ceph job because it definitely takes legitimately longer with all the validations
19:07:32 sean-k-mooney so binding_failed is generally a rare thing so if its happen alto recently that concerning
19:07:39 dansmith and it's still steaming along when it is getting close to its timeout
19:07:51 dansmith I've seen it a number of times lately
19:08:01 sean-k-mooney whats it at currenlty
19:08:10 sean-k-mooney the timeout
19:08:21 sean-k-mooney 2 hours?
19:08:22 dansmith just over two hours
19:08:48 sean-k-mooney we can bump it to 150 mins i guess
19:09:13 sean-k-mooney i like it when test fail and give you the requst id
19:09:41 sean-k-mooney im going to see why that failed on the neutron side and ill let you know if i find anyting
19:12:13 opendevreview Dan Smith proposed openstack/nova master: Bump nova-ceph-multstore timeout https://review.opendev.org/c/openstack/nova/+/882890
19:13:27 sean-k-mooney oh that is in seconds https://zuul-ci.org/docs/zuul/latest/job-content.html#var-zuul.timeout
23:02:53 dansmith here we go
23:05:36 opendevreview Merged openstack/nova master: Use force=True for os-brick disconnect during delete https://review.opendev.org/c/openstack/nova/+/882847
23:05:43 opendevreview Merged openstack/nova master: Enable use of service user token with admin context https://review.opendev.org/c/openstack/nova/+/882852
#openstack-nova - 2023-05-11
00:19:29 dansmith melwitt: oh jeez I didn't realize that wasn't +V from zuul, thanks
00:20:28 melwitt np. it's hard to see through all the fail
00:20:54 dansmith just too many patches really
00:23:21 melwitt huh.. it looks like it's already been running. it was already +W by you.. I didn't think of the +W kicking off the job again /facepalm
00:23:53 dansmith oh I swore it was
00:24:02 dansmith I see your recheck put it back in the arm queue,
00:24:16 dansmith which is only a few minutes old but the other one has been running in regular check for over an hour
00:24:32 dansmith I thought I must have been looking at the wrong on
00:24:32 melwitt yeah. it doesn't restart the normal jobs thankfully
00:24:49 melwitt no, I'm just le dumb
00:25:10 dansmith no, there are too many identical patches :)
00:33:57 dansmith crap, 2023.1 patch is going to fail
00:34:23 dansmith dhcp fail in the guest
00:34:24 dansmith le sigh
00:34:28 melwitt dangit
00:49:17 dansmith top patch is +V waiting for another trip to gateland
00:49:34 dansmith man cripes
00:49:44 dansmith my friggin timeout patch is gonna fail again
00:50:24 melwitt 😑
00:50:24 melwitt 😑
00:50:39 dansmith clearly something ceph related:
00:50:39 dansmith [ 10.258086] Buffer I/O error on dev vda1, logical block 20, lost async page write
00:50:44 dansmith vda is the root disk, which is on ceph
00:51:05 dansmith last time around cinder just stopped creating volumes halfway through (maybe for similar reasons)
00:52:25 dansmith Out of memory: Killed process 55722 (ceph-osd)
00:52:28 dansmith that'll do it like every time
00:52:32 dansmith cripes
00:53:03 melwitt ouch
00:54:42 dansmith ah,
00:54:59 dansmith we're lowering the swap on our fatter job than the ceph job I increased in their repo
00:56:25 opendevreview Dan Smith proposed openstack/nova master: Bump nova-ceph-multstore timeout https://review.opendev.org/c/openstack/nova/+/882890
00:56:33 dansmith melwitt: ^
00:59:06 melwitt I don't understand that sentence 😆
00:59:06 melwitt I don't understand that sentence 😆
00:59:52 melwitt ok nevermind
00:59:55 dansmith wait, in my commit message or
01:00:18 melwitt yeah you explained what I didn't understand in your commit message
01:00:25 dansmith I bumped the ceph job to 8G while jammifiying and cephadmifying it, but we inherit from that and set it down to 4G, which makes no sense because we run even more stuff than they do
01:00:30 dansmith ack
05:41:49 opendevreview Merged openstack/nova stable/2023.1: Use force=True for os-brick disconnect during delete https://review.opendev.org/c/openstack/nova/+/882858

Earlier   Later