Earlier  
Posted Nick Remark
#openstack-nova - 2023-01-13
13:46:08 sean-k-mooney ya we will review it as normal i might have time to take a look later today
13:46:44 sean-k-mooney this code is incldued more or less in the follow up patch so i have glance over it already
13:47:08 sean-k-mooney so likely there will be no feedback but im doing some email stuff right now so dont want to swtich context
13:47:33 kashyap Man, I'm in Gate-hell, these jobs are failing w/ unrelated errors :( - nova-live-migration, nova-multi-cell, nova-ovs-hybrid-plug, nova-grenade-multinode
13:48:04 kashyap I wonder how many of these jobs really deserve to be "voting"
13:51:24 sean-k-mooney all of them
13:51:49 sean-k-mooney they are pretty stabel if there is currently an issue its new and we shoudl investigate that
13:56:20 kashyap Well, that is the "ideal" scenario, assuming there's "unlimited bandwidth" from contributors. I can't possibly keep investigating CI issues all day and week long
13:56:51 sean-k-mooney https://zuul.openstack.org/builds?job_name=nova-live-migration&job_name=nova-ovs-hybrid-plug&job_name=nova-multi-cell&job_name=nova-grenade-multinode&project=openstack%2Fnova&branch=master&skip=0
13:56:52 kashyap I don't know if all of them are deserving, we have to re-evaulate some jobs to see if they're still worth their salt.
13:57:16 sean-k-mooney i do look at those jobs and the ones that you listed are valuable
13:57:21 kashyap What's annoing is all of these jobs passed in the previous run :-(
13:59:37 sean-k-mooney the grenade job failed on test_volume_backed_live_migration
13:59:46 bauzas kashyap: lemme look then, change id ?
14:00:33 sean-k-mooney presumable https://review.opendev.org/c/openstack/nova/+/869587
14:01:02 kashyap bauzas: Thanks for the offer, I have already opened all the failing 4 jobs and seeing what's up one-by-one
14:01:03 sean-k-mooney or https://review.opendev.org/c/openstack/nova/+/869950
14:01:24 bauzas sean-k-mooney: ack, will look
14:01:33 bauzas TGIF
14:01:43 kashyap Previously 'nova-ceph-multistore' was failing due to 'pcp' package unreliability
14:01:55 kashyap bauzas: That's the patch: https://review.opendev.org/c/openstack/nova/+/869950/
14:01:56 sean-k-mooney ya i saw that in one of the other failaure
14:02:00 sean-k-mooney that might be a mirror issue
14:02:12 sean-k-mooney so not related to the job but the cloud it ran on
14:02:32 sean-k-mooney i know that there was issues with one of the ci providers running out os log storage i think during the week
14:02:42 kashyap Hmm, it's a pity that we don't have a way to selectively run the failing jobs (while retaining the older one) - if nothing has changed in a patch
14:02:57 sean-k-mooney we intentially dont because that is dangours
14:03:07 sean-k-mooney its call the green check policy
14:03:15 kashyap I know, the "danger" is introducing accidental regressions
14:03:32 sean-k-mooney the issue is that we use speculative execution in the gate
14:04:06 sean-k-mooney so running one job might mean you end up with each job testing diffent specultivly merged commits
14:04:20 kashyap I wonder what's wrong with this: *if* a patch has not changed from previous iteration, and a job has failed due to unrelated failure, then allow to selectively re-run just that
14:04:35 sean-k-mooney if you made sure the same commits were preseved for the rerun that might be ok but that is not how zuul works
14:04:52 bauzas yup
14:05:04 bauzas and honestly, I prefer this
14:05:06 kashyap Sure, the commit _is_ preserved
14:05:19 sean-k-mooney that not how zuul works
14:05:26 bauzas if we have some jobs that are not ok, we can then make non-voting in case
14:05:36 sean-k-mooney zuul rebases the commit you submit on top of master to test it merged to the curent state of master
14:18:51 sean-k-mooney kashyap: the grenade job failed becasue of rabbitmq
14:18:52 sean-k-mooney : ERROR oslo.messaging._drivers.impl_rabbit [None req-5a9299c4-d386-4e22-8cca-281e29799492 None None] Connection failed: [Errno 111] ECONNREFUSED (retrying in 19.0 seconds): ConnectionRefusedError: [Errno 111] ECONNREFUSED
14:19:04 sean-k-mooney the second compute node did not stack
14:19:26 sean-k-mooney this is the second time i have seen that in two days so that looks like a real issue in devstack
14:19:29 kashyap sean-k-mooney: I see, thank you for looking. Is that an accidental thing?
14:19:44 kashyap Hm, 2nd time 2 days
14:19:44 sean-k-mooney i have only seen that since yesterday
14:19:59 kashyap But it passed just a few hours ago. Seems non-deterministic to me.
14:20:02 sean-k-mooney so i dont know if something changed this week that broke it
14:20:29 sean-k-mooney yes so it might depend on the provider what we shoudl do is check the conoterl and see if its running there
14:21:26 kashyap I see nothing "green" on the controller - https://zuul.opendev.org/t/openstack/build/9c1b4d7afc4c457fa90ae4ceb128dded
14:21:37 kashyap Although it says in _red_, "193 OK, 103 changed, 1 Failure"
14:21:46 kashyap Probably it's just a miscolored thing
14:22:15 sean-k-mooney looks like its running fine on the contoler
14:23:57 sean-k-mooney althouhg nova-compute is failing on the contoler too but that might be after/durign the upgrade
14:24:19 sean-k-mooney n-cpu is running fine initally at least
14:24:37 sean-k-mooney this could jsut be an network connectiy issue between the vms
14:25:41 kashyap Yeah, thought so; the classic "network connectivity" :)
14:25:46 sean-k-mooney grenade is able to run many of the test
14:25:47 sean-k-mooney https://zuul.opendev.org/t/openstack/build/9c1b4d7afc4c457fa90ae4ceb128dded/log/controller/logs/grenade.sh_log.txt
14:26:12 sean-k-mooney im not sure if it is a network connectivy test as it would not have got to grenade if it did not work before the upgrade
14:26:22 sean-k-mooney we run tempest twice for grenade
14:26:29 sean-k-mooney before upgrade and after
14:26:59 kashyap Yeah, your logic makes sense, though - if it is able do upgrade tests via grenade, then the prob is elsewhere
14:27:00 sean-k-mooney tempest.scenario.test_server_multinode.TestServerMultinode.test_schedule_to_all_nodes
14:27:23 sean-k-mooney is failing because the subnode is nolonger able to connect
14:27:29 opendevreview Alexey Stupnikov proposed openstack/nova master: Log some InstanceNotFound exceptions from libvirt https://review.opendev.org/c/openstack/nova/+/863665
14:32:25 kashyap sean-k-mooney: Thanks; I admire your ability to tirelessly look at CI failures
14:33:13 bauzas sean-k-mooney: I briefly looked at kashyap's CI issues and it looked to me transient network issues
14:33:22 sean-k-mooney well i have swaped back to other thing but in generally all active contibutor are exepcted to help look at them
14:33:30 dansmith sean-k-mooney: I saw one of those connection refused to rabbit things yesterday as well
14:33:45 sean-k-mooney bauzas: i would agree but as i said this is the second tim i have seen the rabbit issues
14:33:46 dansmith along with several other failures that has me concerned
14:33:53 sean-k-mooney ya
14:33:55 bauzas lemme check the nodes
14:34:14 sean-k-mooney so i do think that something has regressed in grenade/devstack/job-defintiosn
14:34:52 sean-k-mooney it could be provider related but i think we have seen isseus on more the one provider
14:35:02 dansmith on another note, sean-k-mooney bauzas I hope either of you can look at this soon: https://review.opendev.org/c/openstack/nova/+/866218
14:35:12 bauzas https://f4c3b65ecec7dfeb9b12-92114ee8da6f13c40794db32b7bbd824.ssl.cf5.rackcdn.com/869950/4/check/nova-grenade-multinode/9c1b4d7/zuul-info/inventory.yaml
14:35:16 bauzas ovh
14:35:22 dansmith still waiting on one devstack thing, but it's small, and a dependent placement thing as well
14:35:32 bauzas dansmith: sure I can help
14:35:43 bauzas oh this
14:35:52 bauzas I said I was looking at it
14:35:54 sean-k-mooney dansmith: ya so i merged the placment fixture change last night
14:36:04 dansmith oh sweet thanks
14:36:22 bauzas already half-reviewed gmann's patch
14:36:23 sean-k-mooney bauzas: ack if you dont get to it by your end of day ill try and get to it before i sign off today
14:36:29 kashyap sean-k-mooney: Yes, of course - I look at them (within reason). I was just saying, can't drop all the other responsibilities and tend to it :( Just not enough hours
14:36:30 dansmith bauzas: sorry, but thanks
14:36:47 dansmith kashyap: well, someone has to :/
14:37:08 kashyap dansmith: True, the mythical "someone"; it's the classic "tragedy of the commons" :/
14:37:09 dansmith if we let it get out of control we'll never recover
14:37:35 bauzas sean-k-mooney: gmann: ok, so by now we only check functional tests using legacy RBAC
14:37:37 bauzas https://review.opendev.org/c/openstack/placement/+/869525/3/placement/tests/functional/fixtures/placement.py
14:37:43 bauzas I'm OK with this
14:38:08 bauzas but will we have a FUP modifying our tests for using the new defaults ?
14:38:37 dansmith bauzas: no, the main devstack job runs with old
14:38:42 dansmith and a new devstack job with new
14:39:03 dansmith that's what I had him add, so we could command them still to be old while we transition
14:40:38 sean-k-mooney well bauzas is askign about functional tests

Earlier   Later