Earlier  
Posted Nick Remark
#openstack-nova - 2023-04-28
20:05:35 dansmith so the stack is currently nova->cinder-tempest->tempest->devstack-ceph-plugin ... it's thorny
20:06:17 gouthamr yes, and i extended the chain further :D with -->manila-tempest-plugin --> devstack-plugin-ceph
20:06:30 dansmith going for a world record? :)
20:07:00 dansmith life was so simple when everything fit on a single floppy ..
20:07:05 gouthamr lol :D if someone wasn't looking, they'd think we're building a new fancy feature
20:07:09 dansmith hehe
20:42:28 opendevreview Merged openstack/nova master: Fix a typo in this URL: https://docs.openstack.org/nova/latest/admin/availability-zones.html https://review.opendev.org/c/openstack/nova/+/878797
#openstack-nova - 2023-05-01
12:54:28 opendevreview sean mooney proposed openstack/nova master: [WIP] use alpine instead of cirros https://review.opendev.org/c/openstack/nova/+/881912
13:01:10 opendevreview sean mooney proposed openstack/nova master: [WIP] use alpine instead of cirros https://review.opendev.org/c/openstack/nova/+/881912
13:04:28 opendevreview sean mooney proposed openstack/nova master: [WIP] use alpine instead of cirros https://review.opendev.org/c/openstack/nova/+/881912
13:31:56 dansmith sean-k-mooney: so check this out: https://264a86dd518c5142b8a6-3b04393c7506ab488b4fde073fc22e36.ssl.cf1.rackcdn.com/881764/1/check/cinder-tempest-plugin-lvm-multiattach/d28afe7/testr_results.html
13:32:01 dansmith maybe kashyap too ^
13:32:25 dansmith that almost kinda looks like we attached the disk to the guest and the kernel crashed while probing partitions/filesystems, right?
13:33:10 dansmith in sysfs_add_file_mode_ns()
13:33:49 dansmith or perhaps while creating sysfs entries for the disk perhaps?
16:30:56 dansmith eharney: you around by chance?
16:45:20 eharney dansmith: yes
16:45:43 dansmith eharney: so, I seem to have gotten the nova ceph job down to a single repeatable failure
16:45:56 dansmith it is in the volume extend test, and it fails during cleanup
16:46:25 dansmith it's trying to, I guess, detach the volume from the server before deleting it and then before deleting the volume
16:46:49 dansmith it does *not* fail locally, so I don't think it's something fundamentally broken with new ceph or anything like that
16:46:54 dansmith https://955f32f8268e5d475e65-6c8f4c6e546a0854b4c11cc7c78829ca.ssl.cf5.rackcdn.com/881585/6/check/nova-ceph-multistore/1ecb09a/testr_results.html
16:47:25 eharney dansmith: let me take a look through the logs
16:47:49 eharney detach failing isn't something i'm familiar with (other than it being mentioned here the other day)
16:48:04 dansmith I've just been tracing through the code looking for how this works and it seems to me like everything is working, but perhaps it's just legitimately that the guest doesn't let go of the volume when we detach after the resize happened
16:48:17 dansmith okay, volume detach is without a doubt our most common failure
16:48:34 dansmith eharney: yeah appreciate if you could see if you spot anything in the logs
17:00:13 dansmith I see the guest saw the size change on vdb and also mentions that it's resizing the filesystem on it... that looks like more than just the block device size change,
17:00:41 dansmith so I wonder if it is literally doing a resize2fs on it and that is still happening when we try to issue the detach and that gets us stuck
17:01:02 dansmith because we start the detach less than half a second after the resize happens
17:03:32 eharney dansmith: yeah, i was also just looking down the path of whether the libvirt block resize call is synchronous or not (it looks like nova assumes it is?)
17:04:47 dansmith eharney: synchronous with what? It's synchronous to libvirt from compute, but I don't know that it waits to return until it's delivered to the guest, but definitely stuff like resizing filesystems would happen after that returns
17:05:08 dansmith even still, the test polls for completion of the operation as far as nova is concerned, but the guest stuff would all be async
17:05:34 eharney i see
17:05:56 dansmith I think maybe I should make the test ssh to the guest and see if it's mounted, and perhaps try to unmount it before the test ends or something, to avoid racing with the detach
17:11:10 dansmith the resize happens before cirros is even done with its startup stuff,
17:11:36 dansmith so I wonder on a slow emulated guest we mount all the filesystems we find, which means it's mounted in the guest when we resize, so it does the resize2fs activity automatically
17:11:58 dansmith but on a fast non-nested local run, it finishes startup before we do the attach and thus doesn't end up with it mounted during the resize
17:12:29 eharney is resize2fs etc triggered by qemu-guest-agent?
17:13:08 dansmith could be.. does cirros have the guest agent in it? I assumed not
17:13:22 eharney i don't think so
17:13:25 dansmith on regular systems I've never had resize2fs triggered automatically for me, so I'm not really sure why that would happen
17:13:30 dansmith but it seems like it is here
17:14:00 eharney i would have guessed that resize2fs doesn't happen in these test jobs
17:14:10 dansmith me too
17:14:12 dansmith you see it in the output though right?
17:14:26 eharney no, where is that?
17:14:27 dansmith [ 48.255755] EXT4-fs (vda1): resized filesystem to 259835
17:14:27 dansmith [ 48.156649] EXT4-fs (vda1): resizing filesystem from 25600 to 259835 blocks
17:14:40 dansmith in the guest console dump
17:14:45 dansmith oh damn
17:14:53 dansmith that is vda, nevermind!
17:15:08 dansmith that's cirros resizing its root disk on startup, not the attached volume
17:15:10 eharney ah, right
17:15:21 dansmith mah bad
17:17:19 dansmith so, without a doubt the most common failure in nova jobs is failing to detach volumes
17:17:33 dansmith we've been trying to get a handle on it for a long time,
17:17:42 eharney does it show up on non-rbd volumes?
17:17:48 dansmith yeah
17:18:13 dansmith there is some assertion that if we attach to an instance before it is far enough along during boot, then it might prevent it from being detached later
17:18:26 dansmith I'm slightly skeptical of that, but we've been adding "wait for sshable" checks everywhere
17:19:01 eharney i guess i'm not sure what kind of conditions in the libvirt area would prevent detach from completing
17:19:17 dansmith I just added that for this test recently (merged on friday) which passes normally (and passes locally) but with this rbd job it seems to fail.. it passed the gate on the focal-based rbd, but not on new ceph and jammy
#openstack-nova - 2023-05-02
08:07:20 kashyap dansmith: I was off yesterday (public holiday). Looking at the scrollback now about "tempest.api.compute.admin.test_volumes_negative.VolumesAdminNegativeTest"
09:38:13 opendevreview Elod Illes proposed openstack/nova master: Add nova-tox-functional-py310 to gate jobs https://review.opendev.org/c/openstack/nova/+/881339
11:19:55 sfinucan bauzas: addressed your comments on https://review.opendev.org/c/openstack/python-openstackclient/+/881822 lemme know if they work
11:57:57 sean-k-mooney sfinucan: im ok wiht osc generating keypairs client side
11:58:17 sean-k-mooney althoug it proably should be in its own command but i understand why we might wasnt to proxy it form the nova one
11:59:23 sean-k-mooney sfinucan: can i suggest we do both. proceed with bauzas patch as is and then add your change on top to add client side generation
11:59:39 sean-k-mooney the keypair type would be changing so they are not really interchangable
12:00:05 sean-k-mooney they kind of are but old operating systsms wont supprot ssh-ed25519
12:00:25 sean-k-mooney i.e. centos 7 and newer OSs wont support RSA i.e. centos 9
12:02:42 sean-k-mooney so i see you are proposing taking the approch of always generating htem client side isntaed fo doing this based on the microversion
12:03:21 sean-k-mooney that actully has some benifits in that the behaivor wont be "incorrect" for the new microverion
12:03:50 sean-k-mooney i.e. the client wont be giving the perception of the api generating the key for the latest microversion
12:04:02 sean-k-mooney since now it will be alwasy client side.
12:56:59 sfinucan sean-k-mooney: yeah, I'd rather not have different behaviour depending on the microversion in use
12:58:22 bauzas sfinucan: sean-k-mooney: I'm OK with generating private keys by default in OSC even if the API microversion is older than 2.92
12:58:50 bauzas since the API was able to provide a private key by passing it, I'm ok
13:03:13 sean-k-mooney ok
13:03:19 sean-k-mooney so how do you want to proceed
13:03:24 sean-k-mooney just go with sfinucan patch
13:03:33 sean-k-mooney or merge both?
13:04:42 sfinucan I think it's a case of one or the other
13:06:25 stephenfin sean-k-mooney: if we start doing stuff client-side, there's no reason not to do it for all microversions. If we stick server-side, we're obliged to drop key generation functionality for newer or all microversions
13:11:19 stephenfin sean-k-mooney: bauzas: btw, do you have +2 on OSC now? You should?
13:12:48 bauzas stephenfin: my only concern if we merge your patch is that enduser wouldn' longer see that they need to create their own keypairs
13:13:26 bauzas stephenfin: so if they use openstacksdk directly, they would see then
13:13:33 sean-k-mooney i can check
13:13:46 stephenfin wdym? They don't. We're still doing it for them, only on the client rather than the server
13:13:48 stephenfin Ah
13:14:08 stephenfin I mean, I think that's okay. We do a whole load of helpful things in OSC that don't happen in SDK
13:15:36 stephenfin Like the 'server ssh' command. There's no equivalent for that in SDK, mainly because it's not needed. We could add a docstring in SDK noting that the private key field must be explicitly provided in newer microversions
13:18:33 sean-k-mooney stephenfin: for osc i do not have +2 currently
13:18:54 sean-k-mooney likely the same for sdk i have not actuly looked
13:26:08 bauzas ditto, I'm not osc-core (and not sure I want it :) )
13:27:19 dansmith kashyap: thanks for looking..that looks like a kernel bug to me and so it'd be good to know if it's been fixed as it's not uncommon to see it
13:27:31 stephenfin Ah, okay, it seems I'd misunderstood what 'python-openstackclient-service-core' was being used for. I'd thought it would contain all the service core groups by default. Evidently not.
13:28:27 stephenfin bauzas: https://review.opendev.org/c/openstack/openstacksdk/+/881965
13:29:15 kashyap dansmith: Yeah, it looks like "somehow" the loading of the kernel modules itself is failing: "failed loading these modules: nls_ascii nls_iso8859-1 nls_utf8 ip_tables ahci"

Earlier   Later