| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2023-05-01 | |||
| 17:11:58 | dansmith | but on a fast non-nested local run, it finishes startup before we do the attach and thus doesn't end up with it mounted during the resize | |
| 17:12:29 | eharney | is resize2fs etc triggered by qemu-guest-agent? | |
| 17:13:08 | dansmith | could be.. does cirros have the guest agent in it? I assumed not | |
| 17:13:22 | eharney | i don't think so | |
| 17:13:25 | dansmith | on regular systems I've never had resize2fs triggered automatically for me, so I'm not really sure why that would happen | |
| 17:13:30 | dansmith | but it seems like it is here | |
| 17:14:00 | eharney | i would have guessed that resize2fs doesn't happen in these test jobs | |
| 17:14:10 | dansmith | me too | |
| 17:14:12 | dansmith | you see it in the output though right? | |
| 17:14:26 | eharney | no, where is that? | |
| 17:14:27 | dansmith | [ 48.156649] EXT4-fs (vda1): resizing filesystem from 25600 to 259835 blocks | |
| 17:14:27 | dansmith | [ 48.255755] EXT4-fs (vda1): resized filesystem to 259835 | |
| 17:14:40 | dansmith | in the guest console dump | |
| 17:14:45 | dansmith | oh damn | |
| 17:14:53 | dansmith | that is vda, nevermind! | |
| 17:15:08 | dansmith | that's cirros resizing its root disk on startup, not the attached volume | |
| 17:15:10 | eharney | ah, right | |
| 17:15:21 | dansmith | mah bad | |
| 17:17:19 | dansmith | so, without a doubt the most common failure in nova jobs is failing to detach volumes | |
| 17:17:33 | dansmith | we've been trying to get a handle on it for a long time, | |
| 17:17:42 | eharney | does it show up on non-rbd volumes? | |
| 17:17:48 | dansmith | yeah | |
| 17:18:13 | dansmith | there is some assertion that if we attach to an instance before it is far enough along during boot, then it might prevent it from being detached later | |
| 17:18:26 | dansmith | I'm slightly skeptical of that, but we've been adding "wait for sshable" checks everywhere | |
| 17:19:01 | eharney | i guess i'm not sure what kind of conditions in the libvirt area would prevent detach from completing | |
| 17:19:17 | dansmith | I just added that for this test recently (merged on friday) which passes normally (and passes locally) but with this rbd job it seems to fail.. it passed the gate on the focal-based rbd, but not on new ceph and jammy | |
| #openstack-nova - 2023-05-02 | |||
| 08:07:20 | kashyap | dansmith: I was off yesterday (public holiday). Looking at the scrollback now about "tempest.api.compute.admin.test_volumes_negative.VolumesAdminNegativeTest" | |
| 09:38:13 | opendevreview | Elod Illes proposed openstack/nova master: Add nova-tox-functional-py310 to gate jobs https://review.opendev.org/c/openstack/nova/+/881339 | |
| 11:19:55 | sfinucan | bauzas: addressed your comments on https://review.opendev.org/c/openstack/python-openstackclient/+/881822 lemme know if they work | |
| 11:57:57 | sean-k-mooney | sfinucan: im ok wiht osc generating keypairs client side | |
| 11:58:17 | sean-k-mooney | althoug it proably should be in its own command but i understand why we might wasnt to proxy it form the nova one | |
| 11:59:23 | sean-k-mooney | sfinucan: can i suggest we do both. proceed with bauzas patch as is and then add your change on top to add client side generation | |
| 11:59:39 | sean-k-mooney | the keypair type would be changing so they are not really interchangable | |
| 12:00:05 | sean-k-mooney | they kind of are but old operating systsms wont supprot ssh-ed25519 | |
| 12:00:25 | sean-k-mooney | i.e. centos 7 and newer OSs wont support RSA i.e. centos 9 | |
| 12:02:42 | sean-k-mooney | so i see you are proposing taking the approch of always generating htem client side isntaed fo doing this based on the microversion | |
| 12:03:21 | sean-k-mooney | that actully has some benifits in that the behaivor wont be "incorrect" for the new microverion | |
| 12:03:50 | sean-k-mooney | i.e. the client wont be giving the perception of the api generating the key for the latest microversion | |
| 12:04:02 | sean-k-mooney | since now it will be alwasy client side. | |
| 12:56:59 | sfinucan | sean-k-mooney: yeah, I'd rather not have different behaviour depending on the microversion in use | |
| 12:58:22 | bauzas | sfinucan: sean-k-mooney: I'm OK with generating private keys by default in OSC even if the API microversion is older than 2.92 | |
| 12:58:50 | bauzas | since the API was able to provide a private key by passing it, I'm ok | |
| 13:03:13 | sean-k-mooney | ok | |
| 13:03:19 | sean-k-mooney | so how do you want to proceed | |
| 13:03:24 | sean-k-mooney | just go with sfinucan patch | |
| 13:03:33 | sean-k-mooney | or merge both? | |
| 13:04:42 | sfinucan | I think it's a case of one or the other | |
| 13:06:25 | stephenfin | sean-k-mooney: if we start doing stuff client-side, there's no reason not to do it for all microversions. If we stick server-side, we're obliged to drop key generation functionality for newer or all microversions | |
| 13:11:19 | stephenfin | sean-k-mooney: bauzas: btw, do you have +2 on OSC now? You should? | |
| 13:12:48 | bauzas | stephenfin: my only concern if we merge your patch is that enduser wouldn' longer see that they need to create their own keypairs | |
| 13:13:26 | bauzas | stephenfin: so if they use openstacksdk directly, they would see then | |
| 13:13:33 | sean-k-mooney | i can check | |
| 13:13:46 | stephenfin | wdym? They don't. We're still doing it for them, only on the client rather than the server | |
| 13:13:48 | stephenfin | Ah | |
| 13:14:08 | stephenfin | I mean, I think that's okay. We do a whole load of helpful things in OSC that don't happen in SDK | |
| 13:15:36 | stephenfin | Like the 'server ssh' command. There's no equivalent for that in SDK, mainly because it's not needed. We could add a docstring in SDK noting that the private key field must be explicitly provided in newer microversions | |
| 13:18:33 | sean-k-mooney | stephenfin: for osc i do not have +2 currently | |
| 13:18:54 | sean-k-mooney | likely the same for sdk i have not actuly looked | |
| 13:26:08 | bauzas | ditto, I'm not osc-core (and not sure I want it :) ) | |
| 13:27:19 | dansmith | kashyap: thanks for looking..that looks like a kernel bug to me and so it'd be good to know if it's been fixed as it's not uncommon to see it | |
| 13:27:31 | stephenfin | Ah, okay, it seems I'd misunderstood what 'python-openstackclient-service-core' was being used for. I'd thought it would contain all the service core groups by default. Evidently not. | |
| 13:28:27 | stephenfin | bauzas: https://review.opendev.org/c/openstack/openstacksdk/+/881965 | |
| 13:29:15 | kashyap | dansmith: Yeah, it looks like "somehow" the loading of the kernel modules itself is failing: "failed loading these modules: nls_ascii nls_iso8859-1 nls_utf8 ip_tables ahci" | |
| 13:30:21 | kashyap | Also, I hate to say this, but I wonder if it's CirrOS-specific | |
| 13:30:57 | kashyap | dansmith: I take it that this segfault is consistently reproducible | |
| 13:31:54 | dansmith | kashyap: I can't push a button and make it happen, but I see it fairly often | |
| 13:32:21 | kashyap | dansmith: Noted; I'm trying to wade through the logs to also find the QEMU command-line of this. And the host kernel version | |
| 13:32:28 | dansmith | kashyap: entirely possible that it's cirros related, but since this there is not much going on in userland and this is a kernel crash, I imagine it's not the actual fault of cirros | |
| 13:32:36 | kashyap | Have you got the instance name / UUID, per chance? | |
| 13:35:42 | dansmith | kashyap: if you look right above the top of the console dump you'll see requests tempest was making to watch the status and you should find the uuid there | |
| 13:35:59 | kashyap | Okido, this seems the UUID: e0dc4c98-95e2-4421-a6ba-ae07213914a7 | |
| 13:36:00 | kashyap | Thanks! | |
| 13:40:46 | dansmith | :) | |
| 13:53:25 | dansmith | bauzas: FYI, I've got a few good passing runs on the nova patch with new ceph on jammy, after *much* work on the ceph devstack plugin, cinder tempest plugin, and tempest itself | |
| 13:53:41 | dansmith | so I'm hoping we can start to land those things soon and flip that job over, with your blessing | |
| 13:53:44 | opendevreview | yatin proposed openstack/nova master: Add config option to configure TB cache size https://review.opendev.org/c/openstack/nova/+/868419 | |
| 13:53:52 | dansmith | but let me know if you want to examine results closely | |
| 14:03:03 | bauzas | dansmith: oh that's excellent news, I got distracted from following the series | |
| 14:03:22 | dansmith | gouthamr: I think I need to make the devtack-ceph-plugin dependent on the cinder test plugin, which is dependent on the tempest stack.. are you okay with that? and do you want me to do anything else to the devstack plugin before we merge? or just merge and clean it up after? | |
| 14:03:37 | dansmith | bauzas: not expecting you to follow it, just giving the boss the executive summary :) | |
| 14:03:59 | bauzas | dansmith: don't call me like this, man, I'm just a cat herder :) | |
| 14:04:05 | dansmith | heh | |
| 14:04:51 | bauzas | dansmith: so, what's the plan ? merging the tempest bits first, and then the ceph devstack plugin one ? | |
| 14:05:32 | dansmith | bauzas: well, I was just saying above that I think it needs to be devstack-plugin -> cinder tempest -> tempest (so they merge in reverse order of this) | |
| 14:06:02 | bauzas | sorry missed your line | |
| 14:06:12 | dansmith | and probably change my DNM nova patch to stop commenting out all the jobs to make sure they're all still happy, other than just the ceph one | |
| 14:06:19 | bauzas | yeah | |
| 14:08:43 | bauzas | dansmith: at least that's a good thing to tell in the nova meeting today :) | |
| 14:08:52 | dansmith | yeah | |
| 14:18:57 | sean-k-mooney | dansmith: by the way i pushed this yesterday https://review.opendev.org/c/openstack/nova/+/881912 the image it pulled does not have resize2fs installed | |
| 14:19:09 | sean-k-mooney | so that is why some of the tests failed | |
| 14:19:27 | dansmith | sean-k-mooney: because it can't resize itself to fill the root on boot? | |
| 14:20:08 | sean-k-mooney | yep | |
| 14:20:35 | sean-k-mooney | i can fix that pretty simply i need to look and see if any other packages are missing | |
| 14:21:04 | dansmith | ack, cool | |
| 14:27:40 | sean-k-mooney | in at least some of the case the inablity to resize the root fs prevent the ssh hostkeys from being genreated which prevented ssh access to the vm | |
| 14:28:26 | dansmith | it'll prevent other things too, like writing the timestamp files | |
| 14:28:34 | dansmith | (assuming it's really that constrained) | |
| 14:29:58 | sean-k-mooney | virtual size: 292 MiB (305659904 bytes) | |
| 14:30:00 | sean-k-mooney | disk size: 83 MiB | |