| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-13 | |||
| 13:47:33 | kashyap | Man, I'm in Gate-hell, these jobs are failing w/ unrelated errors :( - nova-live-migration, nova-multi-cell, nova-ovs-hybrid-plug, nova-grenade-multinode | |
| 13:48:04 | kashyap | I wonder how many of these jobs really deserve to be "voting" | |
| 13:51:24 | sean-k-mooney | all of them | |
| 13:51:49 | sean-k-mooney | they are pretty stabel if there is currently an issue its new and we shoudl investigate that | |
| 13:56:20 | kashyap | Well, that is the "ideal" scenario, assuming there's "unlimited bandwidth" from contributors. I can't possibly keep investigating CI issues all day and week long | |
| 13:56:51 | sean-k-mooney | https://zuul.openstack.org/builds?job_name=nova-live-migration&job_name=nova-ovs-hybrid-plug&job_name=nova-multi-cell&job_name=nova-grenade-multinode&project=openstack%2Fnova&branch=master&skip=0 | |
| 13:56:52 | kashyap | I don't know if all of them are deserving, we have to re-evaulate some jobs to see if they're still worth their salt. | |
| 13:57:16 | sean-k-mooney | i do look at those jobs and the ones that you listed are valuable | |
| 13:57:21 | kashyap | What's annoing is all of these jobs passed in the previous run :-( | |
| 13:59:37 | sean-k-mooney | the grenade job failed on test_volume_backed_live_migration | |
| 13:59:46 | bauzas | kashyap: lemme look then, change id ? | |
| 14:00:33 | sean-k-mooney | presumable https://review.opendev.org/c/openstack/nova/+/869587 | |
| 14:01:02 | kashyap | bauzas: Thanks for the offer, I have already opened all the failing 4 jobs and seeing what's up one-by-one | |
| 14:01:03 | sean-k-mooney | or https://review.opendev.org/c/openstack/nova/+/869950 | |
| 14:01:24 | bauzas | sean-k-mooney: ack, will look | |
| 14:01:33 | bauzas | TGIF | |
| 14:01:43 | kashyap | Previously 'nova-ceph-multistore' was failing due to 'pcp' package unreliability | |
| 14:01:55 | kashyap | bauzas: That's the patch: https://review.opendev.org/c/openstack/nova/+/869950/ | |
| 14:01:56 | sean-k-mooney | ya i saw that in one of the other failaure | |
| 14:02:00 | sean-k-mooney | that might be a mirror issue | |
| 14:02:12 | sean-k-mooney | so not related to the job but the cloud it ran on | |
| 14:02:32 | sean-k-mooney | i know that there was issues with one of the ci providers running out os log storage i think during the week | |
| 14:02:42 | kashyap | Hmm, it's a pity that we don't have a way to selectively run the failing jobs (while retaining the older one) - if nothing has changed in a patch | |
| 14:02:57 | sean-k-mooney | we intentially dont because that is dangours | |
| 14:03:07 | sean-k-mooney | its call the green check policy | |
| 14:03:15 | kashyap | I know, the "danger" is introducing accidental regressions | |
| 14:03:32 | sean-k-mooney | the issue is that we use speculative execution in the gate | |
| 14:04:06 | sean-k-mooney | so running one job might mean you end up with each job testing diffent specultivly merged commits | |
| 14:04:20 | kashyap | I wonder what's wrong with this: *if* a patch has not changed from previous iteration, and a job has failed due to unrelated failure, then allow to selectively re-run just that | |
| 14:04:35 | sean-k-mooney | if you made sure the same commits were preseved for the rerun that might be ok but that is not how zuul works | |
| 14:04:52 | bauzas | yup | |
| 14:05:04 | bauzas | and honestly, I prefer this | |
| 14:05:06 | kashyap | Sure, the commit _is_ preserved | |
| 14:05:19 | sean-k-mooney | that not how zuul works | |
| 14:05:26 | bauzas | if we have some jobs that are not ok, we can then make non-voting in case | |
| 14:05:36 | sean-k-mooney | zuul rebases the commit you submit on top of master to test it merged to the curent state of master | |
| 14:18:51 | sean-k-mooney | kashyap: the grenade job failed becasue of rabbitmq | |
| 14:18:52 | sean-k-mooney | : ERROR oslo.messaging._drivers.impl_rabbit [None req-5a9299c4-d386-4e22-8cca-281e29799492 None None] Connection failed: [Errno 111] ECONNREFUSED (retrying in 19.0 seconds): ConnectionRefusedError: [Errno 111] ECONNREFUSED | |
| 14:19:04 | sean-k-mooney | the second compute node did not stack | |
| 14:19:26 | sean-k-mooney | this is the second time i have seen that in two days so that looks like a real issue in devstack | |
| 14:19:29 | kashyap | sean-k-mooney: I see, thank you for looking. Is that an accidental thing? | |
| 14:19:44 | sean-k-mooney | i have only seen that since yesterday | |
| 14:19:44 | kashyap | Hm, 2nd time 2 days | |
| 14:19:59 | kashyap | But it passed just a few hours ago. Seems non-deterministic to me. | |
| 14:20:02 | sean-k-mooney | so i dont know if something changed this week that broke it | |
| 14:20:29 | sean-k-mooney | yes so it might depend on the provider what we shoudl do is check the conoterl and see if its running there | |
| 14:21:26 | kashyap | I see nothing "green" on the controller - https://zuul.opendev.org/t/openstack/build/9c1b4d7afc4c457fa90ae4ceb128dded | |
| 14:21:37 | kashyap | Although it says in _red_, "193 OK, 103 changed, 1 Failure" | |
| 14:21:46 | kashyap | Probably it's just a miscolored thing | |
| 14:22:15 | sean-k-mooney | looks like its running fine on the contoler | |
| 14:23:57 | sean-k-mooney | althouhg nova-compute is failing on the contoler too but that might be after/durign the upgrade | |
| 14:24:19 | sean-k-mooney | n-cpu is running fine initally at least | |
| 14:24:37 | sean-k-mooney | this could jsut be an network connectiy issue between the vms | |
| 14:25:41 | kashyap | Yeah, thought so; the classic "network connectivity" :) | |
| 14:25:46 | sean-k-mooney | grenade is able to run many of the test | |
| 14:25:47 | sean-k-mooney | https://zuul.opendev.org/t/openstack/build/9c1b4d7afc4c457fa90ae4ceb128dded/log/controller/logs/grenade.sh_log.txt | |
| 14:26:12 | sean-k-mooney | im not sure if it is a network connectivy test as it would not have got to grenade if it did not work before the upgrade | |
| 14:26:22 | sean-k-mooney | we run tempest twice for grenade | |
| 14:26:29 | sean-k-mooney | before upgrade and after | |
| 14:26:59 | kashyap | Yeah, your logic makes sense, though - if it is able do upgrade tests via grenade, then the prob is elsewhere | |
| 14:27:00 | sean-k-mooney | tempest.scenario.test_server_multinode.TestServerMultinode.test_schedule_to_all_nodes | |
| 14:27:23 | sean-k-mooney | is failing because the subnode is nolonger able to connect | |
| 14:27:29 | opendevreview | Alexey Stupnikov proposed openstack/nova master: Log some InstanceNotFound exceptions from libvirt https://review.opendev.org/c/openstack/nova/+/863665 | |
| 14:32:25 | kashyap | sean-k-mooney: Thanks; I admire your ability to tirelessly look at CI failures | |
| 14:33:13 | bauzas | sean-k-mooney: I briefly looked at kashyap's CI issues and it looked to me transient network issues | |
| 14:33:22 | sean-k-mooney | well i have swaped back to other thing but in generally all active contibutor are exepcted to help look at them | |
| 14:33:30 | dansmith | sean-k-mooney: I saw one of those connection refused to rabbit things yesterday as well | |
| 14:33:45 | sean-k-mooney | bauzas: i would agree but as i said this is the second tim i have seen the rabbit issues | |
| 14:33:46 | dansmith | along with several other failures that has me concerned | |
| 14:33:53 | sean-k-mooney | ya | |
| 14:33:55 | bauzas | lemme check the nodes | |
| 14:34:14 | sean-k-mooney | so i do think that something has regressed in grenade/devstack/job-defintiosn | |
| 14:34:52 | sean-k-mooney | it could be provider related but i think we have seen isseus on more the one provider | |
| 14:35:02 | dansmith | on another note, sean-k-mooney bauzas I hope either of you can look at this soon: https://review.opendev.org/c/openstack/nova/+/866218 | |
| 14:35:12 | bauzas | https://f4c3b65ecec7dfeb9b12-92114ee8da6f13c40794db32b7bbd824.ssl.cf5.rackcdn.com/869950/4/check/nova-grenade-multinode/9c1b4d7/zuul-info/inventory.yaml | |
| 14:35:16 | bauzas | ovh | |
| 14:35:22 | dansmith | still waiting on one devstack thing, but it's small, and a dependent placement thing as well | |
| 14:35:32 | bauzas | dansmith: sure I can help | |
| 14:35:43 | bauzas | oh this | |
| 14:35:52 | bauzas | I said I was looking at it | |
| 14:35:54 | sean-k-mooney | dansmith: ya so i merged the placment fixture change last night | |
| 14:36:04 | dansmith | oh sweet thanks | |
| 14:36:22 | bauzas | already half-reviewed gmann's patch | |
| 14:36:23 | sean-k-mooney | bauzas: ack if you dont get to it by your end of day ill try and get to it before i sign off today | |
| 14:36:29 | kashyap | sean-k-mooney: Yes, of course - I look at them (within reason). I was just saying, can't drop all the other responsibilities and tend to it :( Just not enough hours | |
| 14:36:30 | dansmith | bauzas: sorry, but thanks | |
| 14:36:47 | dansmith | kashyap: well, someone has to :/ | |
| 14:37:08 | kashyap | dansmith: True, the mythical "someone"; it's the classic "tragedy of the commons" :/ | |
| 14:37:09 | dansmith | if we let it get out of control we'll never recover | |
| 14:37:35 | bauzas | sean-k-mooney: gmann: ok, so by now we only check functional tests using legacy RBAC | |
| 14:37:37 | bauzas | https://review.opendev.org/c/openstack/placement/+/869525/3/placement/tests/functional/fixtures/placement.py | |
| 14:37:43 | bauzas | I'm OK with this | |
| 14:38:08 | bauzas | but will we have a FUP modifying our tests for using the new defaults ? | |
| 14:38:37 | dansmith | bauzas: no, the main devstack job runs with old | |
| 14:38:42 | dansmith | and a new devstack job with new | |
| 14:39:03 | dansmith | that's what I had him add, so we could command them still to be old while we transition | |
| 14:40:38 | sean-k-mooney | well bauzas is askign about functional tests | |
| 14:40:44 | bauzas | dansmith: ok, I understand it so placement won't yet support new defaults, only nova/neutron/cinder blah, right? | |
| 14:40:59 | sean-k-mooney | at some point we shoudl test with new placemnt defaults too | |
| 14:41:05 | bauzas | that's my point | |