| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-13 | |||
| 09:08:54 | kashyap | gibi: Morning, 'grenade-skip-level' and 'nova-ceph-multistore' jobs are failing for me (looks unrelated): https://review.opendev.org/c/openstack/nova/+/869950/ | |
| 09:09:34 | kashyap | One is: | |
| 09:09:36 | kashyap | --- | |
| 09:09:37 | kashyap | dpkg: error processing package pcp (--configure): installed pcp package post-installation script subprocess returned error exit status 1 | |
| 09:09:41 | kashyap | --- | |
| 09:15:59 | gibi | yepp that is unrelated | |
| 09:16:46 | gibi | https://bugs.launchpad.net/devstack/+bug/1943184 | |
| 09:17:10 | kashyap | Ah, thanks for the link | |
| 09:17:27 | kashyap | And the 'nova-ceph-multistore' job seems to crash/segfault Python due to this test: | |
| 09:17:40 | kashyap | tempest.api.compute.admin.test_volume.AttachSCSIVolumeTestJSON.test_attach_scsi_disk_with_config_drive[id-777e468f-17ca-4da4-b93d-b7dbf56c0494] | |
| 09:18:30 | kashyap | gibi: Wow, if 'pcp' has eeb unreliable for that long, I wonder if there's an alternative or if it's necessary at all | |
| 09:19:27 | frickler | it is only for stat collection, so mostly not necessary at all. I was also thinking we had disabled it by default, do you enable dstat in those job(s)? | |
| 09:21:06 | kashyap | frickler: I don't know off-hand if those jobs enable 'dstat', but I assume they do | |
| 10:29:04 | opendevreview | Sahid Orentino Ferdjaoui proposed openstack/nova master: api: extend evacuate instance to support target state https://review.opendev.org/c/openstack/nova/+/858384 | |
| 10:29:04 | opendevreview | Sahid Orentino Ferdjaoui proposed openstack/nova master: compute: enhance compute evacuate instance to support target state https://review.opendev.org/c/openstack/nova/+/858383 | |
| 12:01:54 | opendevreview | Alexey Stupnikov proposed openstack/nova master: Add functional tests to reproduce bug #1994983 https://review.opendev.org/c/openstack/nova/+/863416 | |
| 12:02:08 | opendevreview | Alexey Stupnikov proposed openstack/nova master: Log some InstanceNotFound exceptions from libvirt https://review.opendev.org/c/openstack/nova/+/863665 | |
| 12:06:02 | lajoskatona | Hi nova team, shall I ask about the CLI of migrate? The question is: "is there a chance to change the --wait option to wait for the migrate status instead of the server status in case of openstack server migrate .... --wait?" | |
| 12:08:16 | lajoskatona | The logic is here: https://opendev.org/openstack/python-openstackclient/src/branch/master/openstackclient/compute/v2/server.py#L3016-L3022 and as I saw it was (the login I mean) copy-pasted from novaclient, but for that I can't find why it was decided to wait for the server status instead of the status of the migration | |
| 12:08:54 | lajoskatona | I see reason for both, as even if the migration failed the server can remain on the same host and we are happy as it is active. | |
| 12:09:49 | opendevreview | Alexey Stupnikov proposed openstack/nova stable/zed: Remove deleted projects from flavor access list https://review.opendev.org/c/openstack/nova/+/870053 | |
| 12:10:11 | lajoskatona | But from the other perspective the user would be happy to see in this case that hey your migration failed (without extra check for the status of migration), as it can be misleading that the --wait returns happily but the migration failed | |
| 12:13:18 | sean-k-mooney | lajoskatona: if i recall there is not a good way to find the miration object | |
| 12:14:27 | sean-k-mooney | the migrate and live migrate calls dont retrun the migration uuid if i recall so you would need ot have a hureistic to find it client side | |
| 12:15:25 | sean-k-mooney | something like list the migrations or server events for the instnace and get the last one and hope that is the correct one for the currnt command | |
| 12:16:22 | sean-k-mooney | https://docs.openstack.org/api-ref/compute/?expanded=migrate-server-migrate-action-detail#migrate-server-migrate-action | |
| 12:17:20 | sean-k-mooney | if we had an api change to retrun the migration uuid form that and the live migrate endpoint then it would be easy for the client to wait on the migration status instead | |
| 12:25:28 | lajoskatona | sean-k-mooney: thanks, sounds interesting and true as I start to remember the migration things. I check and play with it to understand fully. | |
| 13:20:38 | pslestang | Hy all, is that because the relation chain is not totally reviewed that I can not merge this patchset https://review.opendev.org/c/openstack/nova/+/867832 or do I miss something else? | |
| 13:36:01 | sean-k-mooney | yes | |
| 13:36:15 | sean-k-mooney | the repoducer is not approved so the fix won be merged | |
| 13:36:39 | sean-k-mooney | when the parent merges the top patch will be merged by zuul | |
| 13:39:56 | pslestang | ok understood, will some of you get some times to approve it? | |
| 13:46:08 | sean-k-mooney | ya we will review it as normal i might have time to take a look later today | |
| 13:46:44 | sean-k-mooney | this code is incldued more or less in the follow up patch so i have glance over it already | |
| 13:47:08 | sean-k-mooney | so likely there will be no feedback but im doing some email stuff right now so dont want to swtich context | |
| 13:47:33 | kashyap | Man, I'm in Gate-hell, these jobs are failing w/ unrelated errors :( - nova-live-migration, nova-multi-cell, nova-ovs-hybrid-plug, nova-grenade-multinode | |
| 13:48:04 | kashyap | I wonder how many of these jobs really deserve to be "voting" | |
| 13:51:24 | sean-k-mooney | all of them | |
| 13:51:49 | sean-k-mooney | they are pretty stabel if there is currently an issue its new and we shoudl investigate that | |
| 13:56:20 | kashyap | Well, that is the "ideal" scenario, assuming there's "unlimited bandwidth" from contributors. I can't possibly keep investigating CI issues all day and week long | |
| 13:56:51 | sean-k-mooney | https://zuul.openstack.org/builds?job_name=nova-live-migration&job_name=nova-ovs-hybrid-plug&job_name=nova-multi-cell&job_name=nova-grenade-multinode&project=openstack%2Fnova&branch=master&skip=0 | |
| 13:56:52 | kashyap | I don't know if all of them are deserving, we have to re-evaulate some jobs to see if they're still worth their salt. | |
| 13:57:16 | sean-k-mooney | i do look at those jobs and the ones that you listed are valuable | |
| 13:57:21 | kashyap | What's annoing is all of these jobs passed in the previous run :-( | |
| 13:59:37 | sean-k-mooney | the grenade job failed on test_volume_backed_live_migration | |
| 13:59:46 | bauzas | kashyap: lemme look then, change id ? | |
| 14:00:33 | sean-k-mooney | presumable https://review.opendev.org/c/openstack/nova/+/869587 | |
| 14:01:02 | kashyap | bauzas: Thanks for the offer, I have already opened all the failing 4 jobs and seeing what's up one-by-one | |
| 14:01:03 | sean-k-mooney | or https://review.opendev.org/c/openstack/nova/+/869950 | |
| 14:01:24 | bauzas | sean-k-mooney: ack, will look | |
| 14:01:33 | bauzas | TGIF | |
| 14:01:43 | kashyap | Previously 'nova-ceph-multistore' was failing due to 'pcp' package unreliability | |
| 14:01:55 | kashyap | bauzas: That's the patch: https://review.opendev.org/c/openstack/nova/+/869950/ | |
| 14:01:56 | sean-k-mooney | ya i saw that in one of the other failaure | |
| 14:02:00 | sean-k-mooney | that might be a mirror issue | |
| 14:02:12 | sean-k-mooney | so not related to the job but the cloud it ran on | |
| 14:02:32 | sean-k-mooney | i know that there was issues with one of the ci providers running out os log storage i think during the week | |
| 14:02:42 | kashyap | Hmm, it's a pity that we don't have a way to selectively run the failing jobs (while retaining the older one) - if nothing has changed in a patch | |
| 14:02:57 | sean-k-mooney | we intentially dont because that is dangours | |
| 14:03:07 | sean-k-mooney | its call the green check policy | |
| 14:03:15 | kashyap | I know, the "danger" is introducing accidental regressions | |
| 14:03:32 | sean-k-mooney | the issue is that we use speculative execution in the gate | |
| 14:04:06 | sean-k-mooney | so running one job might mean you end up with each job testing diffent specultivly merged commits | |
| 14:04:20 | kashyap | I wonder what's wrong with this: *if* a patch has not changed from previous iteration, and a job has failed due to unrelated failure, then allow to selectively re-run just that | |
| 14:04:35 | sean-k-mooney | if you made sure the same commits were preseved for the rerun that might be ok but that is not how zuul works | |
| 14:04:52 | bauzas | yup | |
| 14:05:04 | bauzas | and honestly, I prefer this | |
| 14:05:06 | kashyap | Sure, the commit _is_ preserved | |
| 14:05:19 | sean-k-mooney | that not how zuul works | |
| 14:05:26 | bauzas | if we have some jobs that are not ok, we can then make non-voting in case | |
| 14:05:36 | sean-k-mooney | zuul rebases the commit you submit on top of master to test it merged to the curent state of master | |
| 14:18:51 | sean-k-mooney | kashyap: the grenade job failed becasue of rabbitmq | |
| 14:18:52 | sean-k-mooney | : ERROR oslo.messaging._drivers.impl_rabbit [None req-5a9299c4-d386-4e22-8cca-281e29799492 None None] Connection failed: [Errno 111] ECONNREFUSED (retrying in 19.0 seconds): ConnectionRefusedError: [Errno 111] ECONNREFUSED | |
| 14:19:04 | sean-k-mooney | the second compute node did not stack | |
| 14:19:26 | sean-k-mooney | this is the second time i have seen that in two days so that looks like a real issue in devstack | |
| 14:19:29 | kashyap | sean-k-mooney: I see, thank you for looking. Is that an accidental thing? | |
| 14:19:44 | kashyap | Hm, 2nd time 2 days | |
| 14:19:44 | sean-k-mooney | i have only seen that since yesterday | |
| 14:19:59 | kashyap | But it passed just a few hours ago. Seems non-deterministic to me. | |
| 14:20:02 | sean-k-mooney | so i dont know if something changed this week that broke it | |
| 14:20:29 | sean-k-mooney | yes so it might depend on the provider what we shoudl do is check the conoterl and see if its running there | |
| 14:21:26 | kashyap | I see nothing "green" on the controller - https://zuul.opendev.org/t/openstack/build/9c1b4d7afc4c457fa90ae4ceb128dded | |
| 14:21:37 | kashyap | Although it says in _red_, "193 OK, 103 changed, 1 Failure" | |
| 14:21:46 | kashyap | Probably it's just a miscolored thing | |
| 14:22:15 | sean-k-mooney | looks like its running fine on the contoler | |
| 14:23:57 | sean-k-mooney | althouhg nova-compute is failing on the contoler too but that might be after/durign the upgrade | |
| 14:24:19 | sean-k-mooney | n-cpu is running fine initally at least | |
| 14:24:37 | sean-k-mooney | this could jsut be an network connectiy issue between the vms | |
| 14:25:41 | kashyap | Yeah, thought so; the classic "network connectivity" :) | |
| 14:25:46 | sean-k-mooney | grenade is able to run many of the test | |
| 14:25:47 | sean-k-mooney | https://zuul.opendev.org/t/openstack/build/9c1b4d7afc4c457fa90ae4ceb128dded/log/controller/logs/grenade.sh_log.txt | |
| 14:26:12 | sean-k-mooney | im not sure if it is a network connectivy test as it would not have got to grenade if it did not work before the upgrade | |
| 14:26:22 | sean-k-mooney | we run tempest twice for grenade | |
| 14:26:29 | sean-k-mooney | before upgrade and after | |
| 14:26:59 | kashyap | Yeah, your logic makes sense, though - if it is able do upgrade tests via grenade, then the prob is elsewhere | |
| 14:27:00 | sean-k-mooney | tempest.scenario.test_server_multinode.TestServerMultinode.test_schedule_to_all_nodes | |
| 14:27:23 | sean-k-mooney | is failing because the subnode is nolonger able to connect | |
| 14:27:29 | opendevreview | Alexey Stupnikov proposed openstack/nova master: Log some InstanceNotFound exceptions from libvirt https://review.opendev.org/c/openstack/nova/+/863665 | |
| 14:32:25 | kashyap | sean-k-mooney: Thanks; I admire your ability to tirelessly look at CI failures | |