Earlier  
Posted Nick Remark
#openstack-nova - 2021-11-17
10:20:25 sean-k-mooney[m] stephenfin: so ya i dont think we should have the volume attaments api we have currently and i dont think we should mirror that for the manilla shares going forward
10:22:55 lyarwood sean-k-mooney: I've updated the manila spec FWIW
10:23:08 sean-k-mooney[m] just opened it
10:23:17 sean-k-mooney[m] ill review it this morning
10:23:51 sean-k-mooney[m] i have a doctors appointment in an hour so i might loop back with you later
10:24:04 sean-k-mooney[m] did you see my comment regarding the vm memory
10:25:05 sean-k-mooney[m] oh you going to require file backed memory i almost feel like -2 for that
10:25:11 lyarwood yeah I've suggested going with the simple option for now and queuing the image property work for later on
10:25:21 lyarwood sure go ahead
10:25:29 sean-k-mooney[m] requireing hugepages i could live with
10:25:46 sean-k-mooney[m] file backed memory is not somethign we can schdule on today
10:26:28 sean-k-mooney[m] so there is no way to enforce it so the vm will just not be able to access the shares if it lands on a host without it
10:27:01 lyarwood well the compute would fail the request at that point
10:27:03 lyarwood late on but still
10:27:12 lyarwood we wouldn't have the attachment
10:27:33 sean-k-mooney[m] your expecting a build failure. if its like normal vhost user
10:27:41 sean-k-mooney[m] it will boot but the connectivy wont work
10:27:46 lyarwood well no
10:28:01 lyarwood file backed memory is a configurable on the compute right?
10:28:05 sean-k-mooney[m] yes
10:28:11 lyarwood and with this spec we are only talking about a basic attach share flow
10:28:42 lyarwood if we support shelved it would make this harder to assert but either way
10:28:58 lyarwood during the attach or boot we'd be able to tell if the compute supported file backed or not
10:29:39 sean-k-mooney[m] i guess since this is not boot its not as bad
10:30:00 lyarwood if we support shelved it would be that's more awkward yeah
10:30:40 sean-k-mooney[m] well the other issue is
10:31:17 sean-k-mooney[m] as a normal user you cant tell if file backed memory is used
10:31:37 sean-k-mooney[m] so you dont know if it will work
10:32:23 lyarwood Yeah it's awkward for end users, admins would need file backed host aggregates for this to work I guess
10:32:37 lyarwood but without the image property stuff this is the best we can do in the short term tbh
10:32:56 lyarwood so it's either deliver something this cycle or back it up behind a pile of other work
10:33:00 sean-k-mooney[m] yes but the vm when it was booted did not request “must be able to attached shares” in any way
10:33:38 sean-k-mooney[m] well lets just say use file backed memory or hugepages
10:34:00 sean-k-mooney[m] and i can look at creating the new imge/flavor extra spec in a sperate spec
10:34:20 sean-k-mooney[m] i think it would be a good addtion outside of this feature
10:35:08 sean-k-mooney[m] lyarwood: hugepages also solve the requirement for vhost-user and are user requestable today via flavor or image
10:36:31 sean-k-mooney[m] kashyap by the way do you recall i mention that file backed memory seamed to be not allocating memory form the file
10:36:48 kashyap sean-k-mooney[m]: Very vaguely :)
10:37:01 kashyap sean-k-mooney[m]: Can you refresh my memory, please? Is there a ticket/bug for this?
10:37:04 sean-k-mooney[m] kashyap i have been wondering the last day or two could that be related to tb-cache or something similar
10:37:21 sean-k-mooney[m] no i just deployed it at home to test it
10:37:40 kashyap Interesting. Can you share your guest XML + QEMU command-line to see if I can reproduce it
10:37:43 sean-k-mooney[m] then booted vms and could not over subscibe my ram with OOM
10:38:05 sean-k-mooney[m] well ill have to repoduce it my self in a test envionent
10:38:23 sean-k-mooney[m] but if i do i can share it with you
10:38:50 sean-k-mooney[m] lyarwood: are you setting up a deployment with file backed memory for your manila dev?
10:39:02 lyarwood yeah I plan to
10:39:18 sean-k-mooney[m] ok can you try to reverify the behavior
10:39:49 lyarwood yeah sure
10:40:01 kashyap lyarwood: Hope you're hale and hearty now
10:40:14 kashyap sean-k-mooney[m]: Yeah, that behaviour does sound like tb-cache thing
10:40:30 gibi can I get a second core on this bugfix (bauzas and sean-k-mooney[m] are already positive on it) https://review.opendev.org/c/openstack/nova/+/813419 ?
10:40:30 lyarwood kashyap: yup back to normal now thanks
10:40:50 lyarwood gibi: queued
10:41:00 gibi lyarwood: thanks! I'm glad you are back!
10:41:54 sean-k-mooney[m] basiclly i just tried to boot 6 8G vms on a host with 48G of ram and the 6th one triggered OOM
10:46:10 sean-k-mooney[m] gibi: my +1 dissapeared at some point but its back on it. the config help text is now better then some of our dedicated docs :)
10:46:44 gibi sean-k-mooney[m]: thanks. bauzas pushed me to have proper config doc and even config value validation
10:47:25 sean-k-mooney[m] i proably would have skipped the validation because its easy to miss updating that if we add a new type
10:47:44 sean-k-mooney[m] but we will proably rememeber
10:48:20 sean-k-mooney[m] on the other hand i am seeing a lot of issue related to edgecases with the network vif plugged events
10:48:56 sean-k-mooney[m] which makes me think we need a systematic soluntion to this problem sooner rather then later
10:49:46 sean-k-mooney[m] i might see if i can revie the work to pass the driver form neutron to nova this cycle
10:50:49 sean-k-mooney[m] but without neutron telling us when the event is sent i feel like we will continue to have wack a mole issues
11:11:18 gibi lyarwood: sean-k-mooney[m]: feel free to add me as a reviewer
12:28:00 lyarwood sean-k-mooney: https://review.opendev.org/c/openstack/nova/+/811716 - can you also hit this again today please?
13:06:21 gibi lyarwood: I getting pretty confident that the kernel panic on stable/victoria in nova-live-migration job happens because we are live migrating a guest that is not booted fully up yet. When I added 30 sec sleep before the live migration then the problem dissapeared (5/5 run green)
13:07:05 gibi lyarwood: from the console log I see that that without the sleep tempest would trigger live migration even 10 seconds before the guest fully booted
13:07:12 gibi you can see the run results here https://review.opendev.org/c/openstack/nova/+/817564
13:07:18 lyarwood gibi: kk I was sure I tested my PINGABLE/SSHABLE series against it last week and it still failed
13:07:26 gibi so I think your idea to wait for pingable is a good direction
13:07:44 gibi hm, interesting
13:07:46 lyarwood let me look again
13:08:10 lyarwood https://review.opendev.org/c/openstack/nova/+/817636
13:08:34 lyarwood I want to say that was against https://review.opendev.org/c/openstack/tempest/+/817635/2
13:08:47 lyarwood I'm just cleaning the series up again now and can retest
13:10:36 gibi lyarwood: I don't see kernel panic in the runs of https://review.opendev.org/c/openstack/nova/+/817636, but there are other errors. Let's re-test it and see where we are
13:30:13 opendevreview Lee Yarwood proposed openstack/nova stable/victoria: DNM - Testing volume detach failures https://review.opendev.org/c/openstack/nova/+/817636
13:42:28 sean-k-mooney lyarwood: yes will do
13:43:06 lyarwood thanks
13:48:04 sean-k-mooney ya ok im +1 on that ill get to your spec ater the call im on is over
14:20:30 opendevreview Artom Lifshitz proposed openstack/nova master: DNM: Test token expiration during live migration https://review.opendev.org/c/openstack/nova/+/817778
14:28:45 opendevreview Artom Lifshitz proposed openstack/nova master: DNM: Test token expiration during live migration https://review.opendev.org/c/openstack/nova/+/817778
14:52:19 bauzas folks, in case you don't know, the OpenInfra keynote is starting in 8 mins
15:52:38 gmann gibi: thanks
15:52:54 gibi gmann: sorry I had no time to go back and properly review that today
15:52:55 gibi :/
15:53:13 gmann no worry.
16:13:35 whoami-rajat lyarwood, around?
16:13:50 lyarwood whoami-rajat: hey yeah
16:13:53 lyarwood on a call but can chat
16:13:58 whoami-rajat hey
16:14:00 whoami-rajat ack
16:14:44 whoami-rajat so i don't have any issue with your suggestion of keeping the tried and tested way of nova doing the attachment update, but the team agreed on other flow so don't want to go back and forth
16:15:55 lyarwood yeah appreciate that, I wasn't at PTG so wasn't part of the discussions
16:16:19 lyarwood but as someone maintaining this area more than most I'm still against passing the connector around like this
16:16:34 lyarwood it also keeps the cinder implementation straight forward etc so it's a win win in my view
16:17:21 lyarwood if other nova-specs-cores are against this then they can speak out in the spec
16:18:58 whoami-rajat ack, makes sense to me as the code becomes easier to maintain and debug that way, not sure how much optimization that one less API call does
16:19:45 whoami-rajat i tried discussing the same with other cores in yesterday's nova meeting but we went out of time

Earlier   Later