| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-01-17 | |||
| 09:43:58 | gibi | TZ difference in log? | |
| 09:44:25 | bauzas | gibi: https://paste.opendev.org/show/bI0yvTNy52PzFSQUsGze/ | |
| 09:46:18 | gibi | maybe job-output.txt rendered after the job failed | |
| 09:46:22 | gibi | hm | |
| 09:47:08 | gibi | I would believe the tempest_log over the job-output.txt about the time steps | |
| 09:47:13 | bauzas | me too | |
| 09:47:26 | bauzas | but look, the image eventually was downloaded | |
| 09:47:30 | bauzas | we can see the log | |
| 09:47:44 | bauzas | which means the HTTP call was done | |
| 09:47:55 | gibi | the OOM hit at 22:31:13 based on syslog | |
| 09:48:01 | gibi | that matches the tempest_log timestamp | |
| 09:48:11 | bauzas | good point then | |
| 09:48:21 | bauzas | gibi: I briefly looked at glance logs | |
| 09:48:43 | bauzas | as I said, the image was apparently fully downloaded in 7-ish secs | |
| 09:51:24 | bauzas | oh wait | |
| 09:52:32 | bauzas | gibi: https://paste.opendev.org/show/bLFZGO2MZTYjdRRV3DCM/ | |
| 09:55:07 | bauzas | looks like we were caching the image | |
| 09:55:36 | bauzas | as we got the new path, and then nothing | |
| 09:56:08 | bauzas | and the timings match this time | |
| 10:16:18 | gibi | hm this is interesting, in all the 16 nova-ceph-multistore jobs that failed in the last 10 days the same test case got killed https://paste.opendev.org/show/bmEzF6rFgucUibd4CqTX/ | |
| 10:20:00 | bauzas | gibi: and I guess we'll see the same, which is we want to get the image | |
| 10:32:36 | bauzas | gibi: I'm curious btw., I've seen you using a logsearch tool | |
| 10:33:01 | bauzas | is that a CLI about https://opensearch.logs.openstack.org/ ? | |
| 10:33:24 | gibi | nope, it is https://github.com/gibizer/zuul-log-search | |
| 10:33:44 | gibi | my homebrew tool for grepping zuul logs | |
| 10:50:17 | gibi | I took 6 recent runs and generated the list of test cases run in the killed worker | |
| 10:50:35 | gibi | then I checked for intersection of the set of test cases | |
| 10:50:46 | gibi | and it is only tempest.api.compute.admin.test_volume.AttachSCSIVolumeTestJSON.test_attach_scsi_disk_with_config_drive | |
| 10:50:54 | gibi | the one that got killed | |
| 10:51:37 | gibi | so this points to that single test case as a cause | |
| 10:56:42 | bauzas | gibi: I'm just grabbing another change logs for looking whether the test was also killed while downloading the image | |
| 10:57:19 | gibi | ack | |
| 10:57:51 | bauzas | ah, your tool only downloads a specific file if I use --file | |
| 10:58:02 | bauzas | gibi: can I get all zuul logs from a specific change ? | |
| 10:58:57 | gibi | no, you need to use --file to get a log downloaded. I did it opt-in as I mostly use it to wide search and I wanted to limit the disk and bandwidth usage | |
| 10:59:15 | gibi | feel free to open an issue in the repo to add such option | |
| 10:59:15 | bauzas | k | |
| 10:59:29 | bauzas | I can workaround it for a sec | |
| 11:19:56 | opendevreview | Kashyap Chamarthy proposed openstack/nova master: libvirt: At start-up skip compareCPU() with a workaround https://review.opendev.org/c/openstack/nova/+/870794 | |
| 11:20:26 | kashyap | gibi: When you get a minute, can you have a quick look at the unit test? I know I messed it up slightly but how I'm unclear :/ | |
| 11:36:16 | bauzas | gibi: I tried to look at alot of failing jobs and all of them are indeed failing with the same test | |
| 11:36:35 | bauzas | I tried to find where in https://github.com/openstack/tempest/blob/master/tempest/api/compute/admin/test_volume.py#L76 we have the oomkiller | |
| 11:36:47 | bauzas | but as you said, maybe it's killed after a few seconds | |
| 13:01:15 | sean-k-mooney | bauzas: im going to add a specless bluepint the meeting adgenda and try and implement it before then. we we decided to defer it thats ok but if we agree its trivial enough i would liek to include it in A | |
| 13:03:21 | bauzas | sean-k-mooney: ack | |
| 13:43:21 | zigo | Is there some docs somewhere explaining how to implement an OpenStack wsgi API with keystone auth? | |
| 13:43:49 | zigo | FYI, I already got the db migration with Alembic done ... | |
| 13:43:59 | zigo | (plus oslo_config setup...) | |
| 13:45:45 | zigo | User docs are sometimes lacking info, dev docs are almost inexistant ... :( | |
| 13:55:10 | bauzas | zigo: you are deliberatly left with the choice you want | |
| 13:55:36 | bauzas | you just need to use keystonemiddleware lib | |
| 13:55:58 | bauzas | https://pypi.org/project/keystonemiddleware/ | |
| 13:56:25 | bauzas | https://docs.openstack.org/keystonemiddleware/latest/middlewarearchitecture.html describes the strategies you can choose for Auth'ing | |
| 13:57:01 | bauzas | a recommandation is to use paste for pipelining the WSGI middlewares | |
| 13:58:35 | zigo | Thanks. But there's no code example is shown in the keystonemiddleware's doc. | |
| 13:59:08 | zigo | Like many stuff, I'm stuck with a "look at other project, and attempt cut/past, then see what it does" strategy... | |
| 13:59:48 | bauzas | hah | |
| 13:59:50 | bauzas | that | |
| 14:00:07 | zigo | :) | |
| 14:00:15 | bauzas | yeah, in generall the overall workflow is prescribed, like in https://docs.openstack.org/project-team-guide/index.html | |
| 14:00:48 | bauzas | but beyond this, this is the project's team responsbility to decide how to implement what they want | |
| 14:00:55 | bauzas | like, the WSGI framework they prefer | |
| 14:01:09 | bauzas | or even the WSGI server they'd run with devstack | |
| 14:03:06 | bauzas | zigo: but honestly, the keystonemiddleware plugin isn't that hard to use | |
| 14:03:44 | zigo | I don't think that's the hardest part indeed. I just don't know where to start! :) | |
| 14:03:56 | bauzas | I suppose you just way the regular 'do the auth thing' by keystonmiddleware like in https://docs.openstack.org/keystonemiddleware/latest/middlewarearchitecture.html#authentication-component | |
| 14:04:31 | bauzas | zigo: we have a couple of openstack cookiecutters, if those still exist and are updated | |
| 14:04:57 | bauzas | but yeah, before incepting any code, I'd recommend to formalize your repo structure the openstack way | |
| 14:05:12 | zigo | - api | |
| 14:05:12 | zigo | - oslo.config | |
| 14:05:12 | zigo | - alembic migrations | |
| 14:05:12 | zigo | Yeah, I used it. But it doesn't do: | |
| 14:05:13 | zigo | ... | |
| 14:05:36 | bauzas | :) | |
| 14:05:48 | zigo | Yeah, I'm navigating through many projects to see how they are organized, and I'm trying to pick the best ones. | |
| 14:06:05 | bauzas | if you're asking for a 'Project inception 101 class', I'll make you sad, it doesn't exist :) | |
| 14:06:20 | bauzas | but you can surely bug us if you want guidance | |
| 14:06:35 | bauzas | I guess you know the project team guide ? | |
| 14:06:44 | zigo | Thanks ! :) | |
| 14:06:51 | bauzas | https://docs.openstack.org/project-team-guide/index.html | |
| 14:07:33 | zigo | Well, I know how the community works, gerrit, release management, branches, etc. | |
| 14:07:56 | zigo | I don't think I even need to read this ! :) | |
| 14:07:58 | bauzas | yup, but there is a small but interesting section in that guide https://docs.openstack.org/project-team-guide/technical-guides/index.html | |
| 14:08:04 | opendevreview | Balazs Gibizer proposed openstack/nova master: Use new get_rpc_client API from oslo.messaging https://review.opendev.org/c/openstack/nova/+/869900 | |
| 14:08:33 | bauzas | you also have the API guidelines https://specs.openstack.org/openstack/api-wg/#guidelines | |
| 14:09:30 | bauzas | and then you're left with reading each of the oslo libs docs | |
| 14:09:45 | bauzas | assuming you want RPC | |
| 14:15:42 | zigo | Thanks for all of the links. | |
| 14:15:56 | zigo | I don't think I'll need RPC, but maybe along the way... | |
| 14:29:21 | bauzas | artom: sean-k-mooney: https://review.opendev.org/c/openstack/nova/+/869812 got a weak -1 because I think we need to add an upgrade section in reno | |
| 14:29:57 | bauzas | tl;dr: starting with 2023.1, users could request instance.example.com hostname for their instance, and it would fail | |
| 14:30:10 | bauzas | because of dhcp_domain | |
| 14:31:42 | sean-k-mooney | it wont fail but it will be modifed as currently don | |
| 14:31:52 | sean-k-mooney | but sure lets add that | |
| 14:32:29 | artom | bauzas, sure, OK | |
| 14:33:05 | bauzas | sean-k-mooney: yeah agreed "fail" is too broad | |
| 14:33:19 | bauzas | sean-k-mooney: I mean their instances won't get the hostname they expect | |
| 14:33:31 | bauzas | from the user pov | |
| 14:36:03 | kashyap | gibi: I think for my unit test question in the scroll, it's probably because I accidentally removed a mock. /me tries... | |
| 14:36:26 | gibi | kashyap: sorry, I haven't got back to that yet | |