Earlier  
Posted Nick Remark
#openstack-nova - 2020-06-10
10:54:25 stephenfin agreed
10:54:39 stephenfin if I wasn't changing the function signature, this would be unrelated
10:54:47 stephenfin but because I am, this is a fair change
10:54:54 sean-k-mooney anyway stephen are you going to chekc the hugepage/file backed memeory thing or will i doit
10:55:23 stephenfin sean-k-mooney: If you've time, go for it. I'm going to lunch shortly as I'm in meetings most of this afternoon :(
10:55:28 stephenfin If not, I'll aim to do it before EOD
10:55:46 sean-k-mooney well i want to understand this so that i can figure out how this relates to numa in placmenet
10:56:03 sean-k-mooney so ill try and doit shortly
10:57:06 sean-k-mooney file backed memroy will have to be reported on the ROOT RP
10:57:18 sean-k-mooney since it has no real numa affintiy
10:59:44 gibi I think we can trust the stable cores to see if a code is not py2 compatible
11:00:22 sean-k-mooney the fact that we are not using six anymore will liekly cause more issue then type hints
11:00:34 gibi also if there is such code that valid for py2 and py3 but behaves differently, then during the master review we have to make sure we add enough unit test coverage that will break when the patch is blindy backported to py2
11:00:54 sean-k-mooney yep
11:01:15 gibi in this particular case the interpreter catches the problem
11:01:47 bauzas gibi: sean-k-mooney: what I'm saying is that we somehow opted into thinking more about what we write and what we merge on master now that the ship has sailed
11:01:53 bauzas that's it.
11:01:58 sean-k-mooney list vs iterator/view will be one of the main one for like dict.keys()
11:02:10 stephenfin and unicode/bytes
11:02:20 bauzas previously, we had no brain knobs to do as we were supporting boh
11:02:22 bauzas both*
11:02:22 sean-k-mooney in that case its not as bad
11:02:29 sean-k-mooney as we are sticter in py3
11:02:36 gibi bauzas: sure, we should be aware, but at the same time we should not limit ourselves to write better code on master
11:02:36 sean-k-mooney but yes
11:02:47 bauzas I don't disagree
11:03:15 bauzas I'm just saying we need to ponderate the will of writing nice code with the pragmatism of running such code on different stable releases
11:03:54 bauzas and from someone who sometimes has to backport things down to Newton, I'm thinking about my job for the next 3 years
11:04:13 bauzas if I can ease myself, I'll certainly do
11:06:32 sean-k-mooney stephenfin: so looking at http://paste.openstack.org/show/794558/
11:06:40 sean-k-mooney stephenfin: i think we are good on the numa front
11:09:08 sean-k-mooney stephenfin: this is the relevent virsh output http://paste.openstack.org/show/794560/
11:09:40 sean-k-mooney so libvirt sees the hugepage i have allocated but we are not reporting any hugepages in the host numa_toplogy blob
11:11:50 sean-k-mooney gibi: by the way inital_ram_allocation_ratio still default to 1.5 :( i think we just changed it to 1 in ooo
11:12:29 sean-k-mooney we said we chagned it to 1.0 in the ptg but we did not
11:12:44 gibi sean-k-mooney: yeah, we mentioned 1.0 but I don't remember the context
11:13:03 gibi I do know that the product I helping downstream set it 1.0 directly since Icehose or Mitaka
11:13:19 sean-k-mooney ya and windriver did too
11:13:55 sean-k-mooney maybe we should actully chage the upstream default to 1 but we would have to be very vocal and proably add an upgrade check if we did
11:15:01 sean-k-mooney 16 for our cpus is also nuts as a defalt. anything over about 4 tends to cause issue when the system is loaded
11:15:58 gibi sean-k-mooney: agree about vocalization, but I don't think we should spend too much effort changing this default.
11:15:59 sean-k-mooney chaning initalil_allcoation_raitio actully might be safe as it only matter for new installs
11:16:40 sean-k-mooney gibi: well i think we change it downstream in our product, you change it in yours and windriver and i think mirantis did in theres
11:16:56 sean-k-mooney so if everyone changes the default its proably a bad default
11:17:23 sean-k-mooney gibi: but ya we dont need to spend time on it now
11:17:35 sean-k-mooney i jsut wanted to correct myself
11:17:54 gibi ack, thanks for the correction
11:42:18 aarents Hi nova,
11:42:45 openstack Launchpad bug 1874032 in OpenStack Compute (nova) "nova-compute become stuck when doing IO on busy file system" [Wishlist,In progress] - Assigned to Alexandre arents (aarents)
11:42:45 aarents bauzas: gibi FYI, As per new element given during PTG regarding my issue https://bugs.launchpad.net/nova/+bug/1874032 , I confirm that offload snapshot upload in a native thread fix it
11:42:50 aarents (i'm no more able to get nova-compute hang if file system is busy)
11:45:45 sean-k-mooney stephenfin: actully im going to restart libvirt and test that agian
11:48:34 aarents so I propose some thing here: https://review.opendev.org/#/c/734776/1 dansmith was planning to think about it during PTG
12:15:02 sean-k-mooney aarents: right but didnt we agree not to do that and instead just use a new sepmaphore
12:17:28 sean-k-mooney we have https://github.com/openstack/nova/blob/e954a0569785c941ccb89d31c7b4b847b15b5fc5/doc/source/reference/threading.rst which documents nova threading model
12:18:04 sean-k-mooney there are a few cases where we have to use explcit treading but we avoid it unless absolutly required
12:25:55 aarents sean-k-mooney: thanks for the helpfull link, by new semaphore you mean you mean something limiting max concurrent snapshot ?
12:26:09 sean-k-mooney yes
12:26:25 sean-k-mooney we have a config option to limit the max builds on a host
12:26:45 sean-k-mooney one of the suggestions was to add one for snapshot
12:27:05 sean-k-mooney aarents: just looking at the eventlet docs
12:27:07 sean-k-mooney https://eventlet.net/doc/threading.html#eventlet.tpool.execute
12:27:21 sean-k-mooney so it looks like this will still be a blocking call
12:27:32 sean-k-mooney but it yeild the current greenlet ?
12:28:04 sean-k-mooney e.g. will that internally invoke the function on a native thread
12:28:07 sean-k-mooney yeild
12:28:19 sean-k-mooney then call thread.join at some later point
12:30:18 aarents I understand more like that: coroutine delegate to native thread (block just waiting native thread finish, but in the maintime yeild to other coroutine always)
12:32:35 sean-k-mooney aarents: if that is how it works then it would be ok but the docs are not clear
12:32:58 sean-k-mooney oh i missed the prelude
12:32:59 sean-k-mooney The simplest thing to do with tpool is to execute() a function with it. The function will be run in a random thread in the pool, while the calling coroutine blocks on its completion:
12:34:30 aarents yep
12:34:38 aarents sean-k-mooney: according the nova doc you share, I'm precisely in case "Yielding the thread in long-running tasks" but my issue is: I can't put smart greenthread.sleep(0) call because IO are in glance.client
12:35:08 sean-k-mooney aarents: no that is for internal cases
12:35:12 sean-k-mooney not external io
12:35:22 aarents k
12:35:53 sean-k-mooney if we are computing something expensive we are ment to addd sleeps to yeild
12:36:22 aarents yes, aggree
12:42:50 sean-k-mooney aarents: ok left my toughts in the patch. i would proably be ok with that change if we had the config option to limit the number in parralle
12:43:27 stephenfin bauzas: Last ask for the day, I promise, but could you also take a shot at https://review.opendev.org/#/c/728505/ ? It's a re-approval, if that helps
12:43:58 aarents sean-k-mooney: thanks
12:44:05 stephenfin If you want, will trade for review of your spec :-P
12:52:32 openstackgerrit Merged openstack/nova-specs master: spec update: virtual-persistent-memory https://review.opendev.org/726284
13:05:55 bauzas stephenfin: heh no worries
13:28:22 aarents sean-k-mooney: dansmith thks, I replied https://review.opendev.org/#/c/734776/1
13:32:15 sean-k-mooney aarents: i was thinking more like 4 but 10 might be fine too
13:32:29 sean-k-mooney what is the default build limit
13:32:45 sean-k-mooney ah that is 10
13:33:03 sean-k-mooney ya so 10 would be fine
13:34:15 aarents ok
13:38:46 dansmith 10 seems like a lot to me
13:39:08 dansmith well, depends on where the bottleneck is I guess
13:39:25 sean-k-mooney ya that is why i was thinking 4 but we consider 10 ok for max_concurrent_builds
13:39:42 sean-k-mooney maybe it should default to the value of max_concurrent_builds if its not set
13:39:46 sean-k-mooney to keep them the same
13:40:13 sean-k-mooney we have done that with some of the periodic intervals
13:46:04 aarents sean-k-mooney currently we live without limit by defaut, the new idea is to not exhaust thread pool, in that case 10 is fine but we can align with other parameter why not
13:46:56 sean-k-mooney aarents: well we live without a limit becasue its all in one thread
13:47:24 aarents k
13:48:05 sean-k-mooney update the patch and either have the default depend on the other config value or select a defalut and we can review that then

Earlier   Later