| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2022-09-27 | |||
| 09:48:40 | gibi | bauzas: have a nice one | |
| 09:52:13 | opendevreview | Amit Uniyal proposed openstack/nova stable/yoga: Adds check if blk_dev_info has correct flavor.swap https://review.opendev.org/c/openstack/nova/+/859246 | |
| 09:53:06 | opendevreview | Amit Uniyal proposed openstack/nova stable/wallaby: Adds check if blk_dev_info has correct flavor.swap https://review.opendev.org/c/openstack/nova/+/859247 | |
| 10:28:57 | auniyal__ | gibi, QemuGuestAgentNotEnabled is coming from https://github.com/openstack/nova/blob/512fbdfa9933f2e9b48bcded537ffb394979b24b/nova/virt/libvirt/driver.py#L3223 | |
| 10:30:35 | gibi | auniyal__: could you check image_meta.properties there? is it only hw_qemu_guest_agent missing or os_require_quiesce too? | |
| 10:32:47 | auniyal__ | yes, thats also missing | |
| 10:33:05 | auniyal__ | this is a object, - https://paste.opendev.org/show/bD0rlumubeiQ6Yeowaiw/ | |
| 10:33:34 | auniyal__ | in this - https://paste.opendev.org/show/bV01d0abp2GEVT5Bk4up/ | |
| 10:33:52 | auniyal__ | property is preset in image at line 57 | |
| 10:34:36 | auniyal__ | s/preset/present/ | |
| 10:38:34 | gibi | auniyal__: this point to the direction that we have an issue saving the image properties when we are doing the snapshot | |
| 10:43:57 | auniyal__ | gibi: in here - https://paste.opendev.org/show/bV01d0abp2GEVT5Bk4up/ | |
| 10:44:40 | auniyal__ | before line 88, we have server object, later from this object only we retrieve these image property | |
| 10:44:46 | auniyal__ | *must be | |
| 10:45:22 | auniyal__ | so is there a way, we can check here itself, if propeties gets attached or not, | |
| 10:45:27 | auniyal__ | before calling snapshot | |
| 10:47:51 | auniyal__ | here, https://github.com/openstack/nova/blob/512fbdfa9933f2e9b48bcded537ffb394979b24b/nova/objects/image_meta.py#L114 | |
| 10:49:52 | gibi | auniyal__, bauzas: as far as I see we create some of the image metadata during snapshot here https://github.com/openstack/nova/blob/aad31e6ba489f720f5bdc765c132fd0f059a0329/nova/virt/libvirt/driver.py#L2994 but there is no sign that we intended to copy image properties over to the snapshot. We did copy os_type interestingly but not the rest. The os_type was added there as a bug fix | |
| 10:49:58 | gibi | https://review.opendev.org/c/openstack/nova/+/42877 So now I'm wondering if we want to copy everything there or not | |
| 10:51:32 | gibi | auniyal__: you can try to copy the two image prop your fix need there https://github.com/openstack/nova/blob/aad31e6ba489f720f5bdc765c132fd0f059a0329/nova/virt/libvirt/driver.py#L2994 to see if that helps | |
| 11:29:04 | auniyal__ | gibi, actually this get called https://github.com/openstack/nova/blob/aad31e6ba489f720f5bdc765c132fd0f059a0329/nova/api/openstack/compute/servers.py#L1342 | |
| 11:31:00 | auniyal__ | at this place also - properties are not set | |
| 11:31:01 | auniyal__ | instance.image_meta.properties.get('hw_qemu_guest_agent') | |
| 13:03:37 | bauzas | gibi: good point, that was my guess #2 | |
| 13:33:40 | gibi | ahh this is volume backed snapshot. But then the instance is volume booted, so where qemu_guest_agent is coming from during the original boot from volume? | |
| 13:34:09 | opendevreview | Bence Romsics proposed openstack/nova stable/ussuri: Fix unplugging VIF when migrate/resize VM https://review.opendev.org/c/openstack/nova/+/859433 | |
| 13:34:24 | gibi | I'm confused that we talk about image properties but the instance is booted from volume so there might be no image at all involved | |
| 13:35:35 | opendevreview | Bence Romsics proposed openstack/nova stable/train: Fix unplugging VIF when migrate/resize VM https://review.opendev.org/c/openstack/nova/+/859434 | |
| 13:38:15 | rubasov | hi Nova folks: may I ask for reviews on these two backport series? both merged on master long ago, just taking the backports further: https://review.opendev.org/q/I2c195df5fcf844c0587933b5b5995bdca1a3ebed https://review.opendev.org/q/I3cb39a9ec2c260f422b3c48122b9db512cdd799b | |
| 14:27:58 | Uggla | gibi, bauzas, sorry coming back to this morning discussion. I'm afraid we can have a "deadlock" if we fail to mount a share, we may fail to umount it as well, and we could be stuck with a record. So maybe admin need something to remove this record in the DB ? | |
| 14:29:08 | gibi | Uggla: do you mean the unmount fails *because* we failed to mount the share in the first place? | |
| 14:29:37 | Uggla | I think we could have such kind of situation. | |
| 14:30:00 | Uggla | yes | |
| 14:31:25 | Uggla | and today if a share_mapping is already in the DB, there is an error --> no attempt to mount it again. | |
| 14:31:59 | Uggla | I could change that, to allow to mount it again if it is in error. | |
| 14:35:03 | Uggla | the idea is to discuss how to go out of this if it happens without deleted the db record in the db itself. | |
| 14:35:18 | Uggla | s/deleted/deleting | |
| 14:39:00 | Uggla | I would like also to discuss the behavior starting a vm if some share_mapping are in error. Should I prevent the vm from starting up or start it and do not mount the error shares and warn about it ? | |
| 14:39:46 | gibi | these are good questions | |
| 14:40:44 | Uggla | gibi, fyi also I have changed the behavior from what you saw in the previous code. | |
| 14:40:48 | gibi | so today a ShareMapping in error does not mean that the instance is in ERROR to? | |
| 14:40:51 | gibi | too | |
| 14:41:57 | Uggla | today 1- VM should be stopped, 2- Attach a share --> error --> share_mapping in status error in DB. | |
| 14:43:26 | Uggla | So here if the user starts the VM, I'm in favor of starting it and do not mount the faulty shares. | |
| 14:43:31 | gibi | ack | |
| 14:44:14 | Uggla | User can still see the reason why the shares are not mounted because the share_mapping has a error status. | |
| 14:44:25 | Uggla | *have | |
| 14:44:33 | Uggla | *an | |
| 14:45:41 | gibi | so at 2- nova tries to mount the share to the host the VM is on? | |
| 14:46:11 | gibi | and that mount can fail | |
| 14:46:51 | gibi | and 2- is sync so it does two thing i) puts the ShareMapping in ERROR ii) sends back http 500 to the user | |
| 14:48:13 | Uggla | yep except vm is off. | |
| 14:49:01 | gibi | and I guess if the VM is not on any host (never scheduled or it is shelved_offloaded) then we don't allow to attach a share | |
| 14:49:45 | Uggla | yep with this version the vm should exist and be powered off (shelve is not allowed.) | |
| 14:49:54 | gibi | cool | |
| 14:51:46 | gibi | so the only way to re-try the mount of the failes ShareMapping is to detach and then attach the share again. But you are worried what if detach fails at unmount. I think that is fine. If the system is in that bad shape then the admin needs to intervene | |
| 14:52:54 | gibi | i'm not sure what is the exact low lever scenario when the mount fails and leaves the host in a state where unmount is not possible, but I think we don't have to automate a fix for that right now | |
| 14:53:25 | gibi | if it turns out there a common case where mount then unmount fails, then probably there will be a common solution for that, and then we can add that common solution to our unmount codepath to try | |
| 14:54:02 | Uggla | exactly, but I'm afraid that the only solution for admin will be to remove the share in the DB itself. Not really convenient. | |
| 14:54:34 | gibi | no, the admin can fix the host in a way that the next detach share call that tries the unmount will succeed | |
| 14:54:44 | gibi | s/can/should be able to/ | |
| 14:55:17 | gibi | so we need to allow detach share to be called on an ShareMapping in error | |
| 14:55:23 | gibi | to re-try the unmount | |
| 14:55:51 | Uggla | gibi, yes we can detach a share in error. | |
| 14:56:15 | gibi | the that is our release valve, they can hit detach repeatedly until it succeeds :) | |
| 14:56:37 | gibi | I mean admin can tweak the host and re-try the unmount via detach :) | |
| 14:57:19 | Uggla | yep, probable resolution will be to mount the share manually to have a "correct" state --> then the detach should umount properly. | |
| 14:59:41 | Uggla | ok I'll go in this way, we could still change later. thx gibi . | |
| 15:00:36 | gibi | Uggla: also we can implement unmount in a way that if it sees no mounted share then it returns OK so the admin can just clean up a bad mount then let unmount see that no mount left to remove | |
| 15:04:36 | Uggla | gibi, you are right ! I need to better check what happen in that case with the current code as I reused it. Maybe it is already handled properly like this. | |
| 15:05:04 | gibi | ack | |
| 15:20:53 | bauzas | gibi: Uggla: sorry was at the school with kid | |
| 15:27:37 | Uggla | bauzas, no worries. | |
| 15:28:47 | bauzas | Uggla: so yeah, if we have a state for the share, would be nice | |
| 15:29:42 | Uggla | bauzas, yep we have the status | |
| 15:30:10 | Uggla | so error are "tracked" in the db. | |
| 15:30:15 | Uggla | *errors | |
| 15:33:37 | bauzas | cool | |
| 15:50:10 | bauzas | reminder : nova meeting in 10 mins | |
| 16:00:18 | opendevmeet | The meeting name has been set to 'nova' | |
| 16:00:18 | opendevmeet | Useful Commands: #action #agreed #help #info #idea #link #topic #startvote. | |
| 16:00:18 | opendevmeet | Meeting started Tue Sep 27 16:00:18 2022 UTC and is due to finish in 60 minutes. The chair is bauzas. Information about MeetBot at http://wiki.debian.org/MeetBot. | |
| 16:00:18 | bauzas | #startmeeting nova | |
| 16:00:43 | bauzas | hello, hola, hi, bonjour | |
| 16:00:48 | justas_napa | Hi | |
| 16:01:01 | justas_napa | I'd like to propose a topic | |
| 16:01:07 | elodilles | o/ | |
| 16:01:41 | justas_napa | I'd like to discuss the actions needed to add support for Napatech SmartNIC in Nova | |
| 16:01:47 | bauzas | #link https://wiki.openstack.org/wiki/Meetings/Nova#Agenda_for_next_meeting | |
| 16:01:56 | bauzas | justas_napa: add your topic at the end of ^ | |
| 16:02:18 | bauzas | and we'll discuss it during the open discussion topic | |
| 16:03:05 | gibi | o/ | |
| 16:03:08 | bauzas | ok, let's start and people will join | |
| 16:03:26 | bauzas | #topic Bugs (stuck/critical) | |
| 16:03:32 | bauzas | #info No Critical bug | |
| 16:03:38 | bauzas | #link https://bugs.launchpad.net/nova/+bugs?search=Search&field.status=New 5 new untriaged bugs (+0 since the last meeting) | |
| 16:03:52 | bauzas | I looked at them and I'd like to discuss about two bug reports | |
| 16:04:06 | bauzas | but let's do this after the other pointds | |
| 16:04:15 | bauzas | #link https://storyboard.openstack.org/#!/project/openstack/placement 26 open stories (+0 since the last meeting) in Storyboard for Placement | |
| 16:04:21 | bauzas | #info Add yourself in the team bug roster if you want to help https://etherpad.opendev.org/p/nova-bug-triage-roster | |