Earlier  
Posted Nick Remark
#openstack-nova - 2022-09-27
14:30:00 Uggla yes
14:31:25 Uggla and today if a share_mapping is already in the DB, there is an error --> no attempt to mount it again.
14:31:59 Uggla I could change that, to allow to mount it again if it is in error.
14:35:03 Uggla the idea is to discuss how to go out of this if it happens without deleted the db record in the db itself.
14:35:18 Uggla s/deleted/deleting
14:39:00 Uggla I would like also to discuss the behavior starting a vm if some share_mapping are in error. Should I prevent the vm from starting up or start it and do not mount the error shares and warn about it ?
14:39:46 gibi these are good questions
14:40:44 Uggla gibi, fyi also I have changed the behavior from what you saw in the previous code.
14:40:48 gibi so today a ShareMapping in error does not mean that the instance is in ERROR to?
14:40:51 gibi too
14:41:57 Uggla today 1- VM should be stopped, 2- Attach a share --> error --> share_mapping in status error in DB.
14:43:26 Uggla So here if the user starts the VM, I'm in favor of starting it and do not mount the faulty shares.
14:43:31 gibi ack
14:44:14 Uggla User can still see the reason why the shares are not mounted because the share_mapping has a error status.
14:44:25 Uggla *have
14:44:33 Uggla *an
14:45:41 gibi so at 2- nova tries to mount the share to the host the VM is on?
14:46:11 gibi and that mount can fail
14:46:51 gibi and 2- is sync so it does two thing i) puts the ShareMapping in ERROR ii) sends back http 500 to the user
14:48:13 Uggla yep except vm is off.
14:49:01 gibi and I guess if the VM is not on any host (never scheduled or it is shelved_offloaded) then we don't allow to attach a share
14:49:45 Uggla yep with this version the vm should exist and be powered off (shelve is not allowed.)
14:49:54 gibi cool
14:51:46 gibi so the only way to re-try the mount of the failes ShareMapping is to detach and then attach the share again. But you are worried what if detach fails at unmount. I think that is fine. If the system is in that bad shape then the admin needs to intervene
14:52:54 gibi i'm not sure what is the exact low lever scenario when the mount fails and leaves the host in a state where unmount is not possible, but I think we don't have to automate a fix for that right now
14:53:25 gibi if it turns out there a common case where mount then unmount fails, then probably there will be a common solution for that, and then we can add that common solution to our unmount codepath to try
14:54:02 Uggla exactly, but I'm afraid that the only solution for admin will be to remove the share in the DB itself. Not really convenient.
14:54:34 gibi no, the admin can fix the host in a way that the next detach share call that tries the unmount will succeed
14:54:44 gibi s/can/should be able to/
14:55:17 gibi so we need to allow detach share to be called on an ShareMapping in error
14:55:23 gibi to re-try the unmount
14:55:51 Uggla gibi, yes we can detach a share in error.
14:56:15 gibi the that is our release valve, they can hit detach repeatedly until it succeeds :)
14:56:37 gibi I mean admin can tweak the host and re-try the unmount via detach :)
14:57:19 Uggla yep, probable resolution will be to mount the share manually to have a "correct" state --> then the detach should umount properly.
14:59:41 Uggla ok I'll go in this way, we could still change later. thx gibi .
15:00:36 gibi Uggla: also we can implement unmount in a way that if it sees no mounted share then it returns OK so the admin can just clean up a bad mount then let unmount see that no mount left to remove
15:04:36 Uggla gibi, you are right ! I need to better check what happen in that case with the current code as I reused it. Maybe it is already handled properly like this.
15:05:04 gibi ack
15:20:53 bauzas gibi: Uggla: sorry was at the school with kid
15:27:37 Uggla bauzas, no worries.
15:28:47 bauzas Uggla: so yeah, if we have a state for the share, would be nice
15:29:42 Uggla bauzas, yep we have the status
15:30:10 Uggla so error are "tracked" in the db.
15:30:15 Uggla *errors
15:33:37 bauzas cool
15:50:10 bauzas reminder : nova meeting in 10 mins
16:00:18 opendevmeet The meeting name has been set to 'nova'
16:00:18 opendevmeet Useful Commands: #action #agreed #help #info #idea #link #topic #startvote.
16:00:18 opendevmeet Meeting started Tue Sep 27 16:00:18 2022 UTC and is due to finish in 60 minutes. The chair is bauzas. Information about MeetBot at http://wiki.debian.org/MeetBot.
16:00:18 bauzas #startmeeting nova
16:00:43 bauzas hello, hola, hi, bonjour
16:00:48 justas_napa Hi
16:01:01 justas_napa I'd like to propose a topic
16:01:07 elodilles o/
16:01:41 justas_napa I'd like to discuss the actions needed to add support for Napatech SmartNIC in Nova
16:01:47 bauzas #link https://wiki.openstack.org/wiki/Meetings/Nova#Agenda_for_next_meeting
16:01:56 bauzas justas_napa: add your topic at the end of ^
16:02:18 bauzas and we'll discuss it during the open discussion topic
16:03:05 gibi o/
16:03:08 bauzas ok, let's start and people will join
16:03:26 bauzas #topic Bugs (stuck/critical)
16:03:32 bauzas #info No Critical bug
16:03:38 bauzas #link https://bugs.launchpad.net/nova/+bugs?search=Search&field.status=New 5 new untriaged bugs (+0 since the last meeting)
16:03:52 bauzas I looked at them and I'd like to discuss about two bug reports
16:04:06 bauzas but let's do this after the other pointds
16:04:15 bauzas #link https://storyboard.openstack.org/#!/project/openstack/placement 26 open stories (+0 since the last meeting) in Storyboard for Placement
16:04:21 bauzas #info Add yourself in the team bug roster if you want to help https://etherpad.opendev.org/p/nova-bug-triage-roster
16:04:36 bauzas so, about the bugs,
16:04:58 bauzas #link https://bugs.launchpad.net/nova/+bug/1955035
16:05:34 bauzas looks to me an incomplete bug report as we need to verify whether it's also a problem for devstack
16:05:44 bauzas thoughts ?
16:07:02 gibi I haven't looked at it
16:07:34 bauzas I have another bug report
16:07:51 bauzas #link https://bugs.launchpad.net/nova/+bug/1981562
16:08:15 bauzas I think we could also ask the reporter to verify the comment provided by melwitt
16:08:54 gibi sure the later can be put in Incomplete with a comment about that ^^
16:09:31 bauzas OK, I'll do it
16:09:41 bauzas anyway, that's it for me
16:10:14 Uggla o/
16:10:14 bauzas elodilles: can then you got the bug baton for this week ?
16:10:54 elodilles bauzas: well, i meant last week that the *following* weeks will be a bit busy to me o:)
16:11:36 elodilles i mean, hopefully it will be a calm period, but still, we have the release next week ;)
16:11:51 elodilles so i'd rather skip one more week
16:12:35 elodilles sorry for not writing it clearly last week :S
16:12:37 bauzas elodilles: okay, then I can continue to look about the bug reports next week
16:13:31 bauzas #info bug baton is continued to be used by bauzas for this week
16:13:46 elodilles thanks & sorry :S
16:13:51 bauzas any other bugs, folks ?
16:13:54 bauzas elodilles: np at all
16:15:11 bauzas moving on
16:15:21 bauzas #topic Gate status
16:15:36 bauzas #link https://bugs.launchpad.net/nova/+bugs?field.tag=gate-failure Nova gate bugs
16:15:41 bauzas #link https://zuul.openstack.org/builds?project=openstack%2Fplacement&pipeline=periodic-weekly Placement periodic job status
16:15:46 bauzas #link https://zuul.openstack.org/builds?job_name=tempest-integrated-compute-centos-9-stream&project=openstack%2Fnova&pipeline=periodic-weekly Centos 9 Stream periodic job status
16:15:51 bauzas #link https://zuul.opendev.org/t/openstack/builds?job_name=nova-emulation&pipeline=periodic-weekly&skip=0 Emulation periodic job runs
16:16:02 bauzas all the runs are good ^
16:16:10 bauzas #info Please look at the gate failures and file a bug report with the gate-failure tag.
16:16:18 bauzas as a reminder,
16:16:20 bauzas #info STOP DOING BLIND RECHECKS aka. 'recheck' https://docs.openstack.org/project-team-guide/testing.html#how-to-handle-test-failures

Earlier   Later