<div dir="ltr"><div dir="ltr">&gt; so lustre_inode_cache is the real culprit when signal_cache appears to<br>&gt;  be large.<br>&gt; This cache is slaved on the common inode cache, so there should be one<br>&gt; entry for each lustre inode that is in memory.<br>&gt; These inodes should get pruned when they&#39;ve been inactive for a while.</div><div dir="ltr"><br></div><div>What triggers the prunning?</div><div><br></div><div>&gt;If you look in /proc/sys/fs/inode-nr  there should be two numbers:<br>&gt;  The first is the total number of in-memory inodes for all filesystems.<br>&gt;  The second is the number of &quot;unused&quot; inodes.<br>
&gt;<br>&gt;  When you write &quot;3&quot; to drop_caches, the second number should drop down to<br>
&gt; nearly zero (I get 95 on my desktop, down from 6524).</div><div><br></div><div>Ok, that is useful to know but echoing 3 to drop_cache or generating memory pressure</div><div>clears most of the signal_cache (inode) as well as other lustre objects, so this is working fine.</div><div><br></div><div>The issue that remains is that they are marked as SUnreclaim vs SReclaimable. </div><div>So i do not think there is a memory leak per se. <br></div><div><br></div><div>Regards.</div><div>Jacek Tomaka<br></div><br><div class="gmail_quote"><div dir="ltr" class="gmail_attr">On Mon, Apr 29, 2019 at 1:39 PM NeilBrown &lt;<a href="mailto:neilb@suse.com">neilb@suse.com</a>&gt; wrote:<br></div><blockquote class="gmail_quote" style="margin:0px 0px 0px 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex"><br>
Thanks Jacek,<br>
 so lustre_inode_cache is the real culprit when signal_cache appears to<br>
 be large.<br>
 This cache is slaved on the common inode cache, so there should be one<br>
 entry for each lustre inode that is in memory.<br>
 These inodes should get pruned when they&#39;ve been inactive for a while.<br>
<br>
 If you look in /proc/sys/fs/inode-nr  there should be two numbers:<br>
  The first is the total number of in-memory inodes for all filesystems.<br>
  The second is the number of &quot;unused&quot; inodes.<br>
<br>
 When you write &quot;3&quot; to drop_caches, the second number should drop down to<br>
 nearly zero (I get 95 on my desktop, down from 6524).<br>
<br>
 When signal_cache stays large even after the drop_caches, it suggest<br>
 that there are lots of lustre inodes that are thought to be still<br>
 active.   I&#39;d have to do a bit of digging to understand what that means,<br>
 and a lot more to work out why lustre is holding on to inodes longer<br>
 than you would expect (if that actually is the case).<br>
<br>
 If an inode still has cached data pages attached that cannot easily be<br>
 removed, it will not be purged even if it is unused.<br>
 So if you see the &quot;unused&quot; number remaining high even after a<br>
 &quot;drop_caches&quot;, that might mean that lustre isn&#39;t letting go of cache<br>
 pages for some reason.<br>
<br>
NeilBrown<br>
<br>
<br>
<br>
On Mon, Apr 29 2019, Jacek Tomaka wrote:<br>
<br>
&gt; Wow, Thanks Nathan and NeilBrown.<br>
&gt; It is great to learn about slub merging. It is awesome to have a<br>
&gt; reproducer.<br>
&gt; I am yet to trigger my original problem with slurm_nomerge but<br>
&gt; slabinfo tool (in kernel sources) can actually show merged caches:<br>
&gt; kernel/3.10.0-693.5.2.el7/tools/slabinfo  -a<br>
&gt;<br>
&gt; :t-0000112   &lt;- sysfs_dir_cache kernfs_node_cache blkdev_integrity<br>
&gt; task_delay_info<br>
&gt; :t-0000144   &lt;- flow_cache cl_env_kmem<br>
&gt; :t-0000160   &lt;- sigqueue lov_object_kmem<br>
&gt; :t-0000168   &lt;- lovsub_object_kmem osc_extent_kmem<br>
&gt; :t-0000176   &lt;- vvp_object_kmem nfsd4_stateids<br>
&gt; :t-0000192   &lt;- ldlm_resources kiocb cred_jar inet_peer_cache key_jar<br>
&gt; file_lock_cache kmalloc-192 dmaengine-unmap-16 bio_integrity_payload<br>
&gt; :t-0000216   &lt;- vvp_session_kmem vm_area_struct<br>
&gt; :t-0000256   &lt;- biovec-16 ip_dst_cache bio-0 ll_file_data kmalloc-256<br>
&gt; sgpool-8 filp request_sock_TCP rpc_tasks request_sock_TCPv6<br>
&gt; skbuff_head_cache pool_workqueue lov_thread_kmem<br>
&gt; :t-0000264   &lt;- osc_lock_kmem numa_policy<br>
&gt; :t-0000328   &lt;- osc_session_kmem taskstats<br>
&gt; :t-0000576   &lt;- kioctx xfrm_dst_cache vvp_thread_kmem<br>
&gt; :t-0001152   &lt;- signal_cache lustre_inode_cache<br>
&gt;<br>
&gt; It is not on a machine that had the problem i described before but the<br>
&gt; kernel version is the same so I am assuming the cache merges are the same.<br>
&gt;<br>
&gt; Looks like signal_cache points to lustre_inode_cache.<br>
&gt; Regards.<br>
&gt; Jacek Tomaka<br>
&gt;<br>
&gt;<br>
&gt; On Thu, Apr 25, 2019 at 7:42 AM NeilBrown &lt;<a href="mailto:neilb@suse.com" target="_blank">neilb@suse.com</a>&gt; wrote:<br>
&gt;<br>
&gt;&gt;<br>
&gt;&gt; Hi,<br>
&gt;&gt;  you seem to be able to reproduce this fairly easily.<br>
&gt;&gt;  If so, could you please boot with the &quot;slub_nomerge&quot; kernel parameter<br>
&gt;&gt;  and then reproduce the (apparent) memory leak.<br>
&gt;&gt;  I&#39;m hoping that this will show some other slab that is actually using<br>
&gt;&gt;  the memory - a slab with very similar object-size to signal_cache that<br>
&gt;&gt;  is, by default, being merged with signal_cache.<br>
&gt;&gt;<br>
&gt;&gt; Thanks,<br>
&gt;&gt; NeilBrown<br>
&gt;&gt;<br>
&gt;&gt;<br>
&gt;&gt; On Wed, Apr 24 2019, Nathan Dauchy - NOAA Affiliate wrote:<br>
&gt;&gt;<br>
&gt;&gt; &gt; On Mon, Apr 15, 2019 at 9:18 PM Jacek Tomaka &lt;<a href="mailto:jacekt@dug.com" target="_blank">jacekt@dug.com</a>&gt; wrote:<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt;<br>
&gt;&gt; &gt;&gt; &gt;signal_cache should have one entry for each process (or thread-group).<br>
&gt;&gt; &gt;&gt;<br>
&gt;&gt; &gt;&gt; That is what i thought as well, looking at the kernel source,<br>
&gt;&gt; allocations<br>
&gt;&gt; &gt;&gt; from<br>
&gt;&gt; &gt;&gt; signal_cache happen only during fork.<br>
&gt;&gt; &gt;&gt;<br>
&gt;&gt; &gt;&gt;<br>
&gt;&gt; &gt; I was recently chasing an issue with clients suffering from low memory<br>
&gt;&gt; and<br>
&gt;&gt; &gt; saw that &quot;signal_cache&quot; was a major player.  But the workload on those<br>
&gt;&gt; &gt; clients was not doing a lot of forking.  (and I don&#39;t *think* threading<br>
&gt;&gt; &gt; either)  Rather it was a LOT of metadata read operations.<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; You can see the symptoms by a simple &quot;du&quot; on a Lustre file system:<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; # grep signal_cache /proc/slabinfo<br>
&gt;&gt; &gt; signal_cache         967   1092   1152   28    8 : tunables    0    0<br>
&gt;&gt; 0<br>
&gt;&gt; &gt; : slabdata     39     39      0<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; # du -s /mnt/lfs1/projects/foo<br>
&gt;&gt; &gt; 339744908 /mnt/lfs1/projects/foo<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; # grep signal_cache /proc/slabinfo<br>
&gt;&gt; &gt; signal_cache      164724 164724   1152   28    8 : tunables    0    0<br>
&gt;&gt; 0<br>
&gt;&gt; &gt; : slabdata   5883   5883      0<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; # slabtop -s c -o | head -n 20<br>
&gt;&gt; &gt;  Active / Total Objects (% used)    : 3660791 / 3662863 (99.9%)<br>
&gt;&gt; &gt;  Active / Total Slabs (% used)      : 93019 / 93019 (100.0%)<br>
&gt;&gt; &gt;  Active / Total Caches (% used)     : 72 / 107 (67.3%)<br>
&gt;&gt; &gt;  Active / Total Size (% used)       : 836474.91K / 837502.16K (99.9%)<br>
&gt;&gt; &gt;  Minimum / Average / Maximum Object : 0.01K / 0.23K / 12.75K<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt;   OBJS ACTIVE  USE OBJ SIZE  SLABS OBJ/SLAB CACHE SIZE NAME<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; 164724 164724 100%    1.12K   5883       28    188256K signal_cache<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; 331712 331712 100%    0.50K  10366       32    165856K ldlm_locks<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; 656896 656896 100%    0.12K  20528       32     82112K kmalloc-128<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; 340200 339971  99%    0.19K   8100       42     64800K kmalloc-192<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; 162838 162838 100%    0.30K   6263       26     50104K osc_object_kmem<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; 744192 744192 100%    0.06K  11628       64     46512K kmalloc-64<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; 205128 205128 100%    0.19K   4884       42     39072K dentry<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt;   4268   4256  99%    8.00K   1067        4     34144K kmalloc-8192<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; 162978 162978 100%    0.17K   3543       46     28344K vvp_object_kmem<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; 162792 162792 100%    0.16K   6783       24     27132K<br>
&gt;&gt; kvm_mmu_page_header<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; 162825 162825 100%    0.16K   6513       25     26052K sigqueue<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt;  16368  16368 100%    1.02K    528       31     16896K nfs_inode_cache<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt;  20385  20385 100%    0.58K    755       27     12080K inode_cache<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; Repeat that for more (and bigger) directories and slab cache added up to<br>
&gt;&gt; &gt; more than half the memory on this 24GB node.<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; This is with CentOS-7.6 and lustre-2.10.5_ddn6.<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; I worked around the problem by tackling the &quot;ldlm_locks&quot; memory usage<br>
&gt;&gt; with:<br>
&gt;&gt; &gt; # lctl set_param ldlm.namespaces.lfs*.lru_max_age=10000<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; ...but I did not find a way to reduce the &quot;signal_cache&quot;.<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; Regards,<br>
&gt;&gt; &gt; Nathan<br>
&gt;&gt;<br>
&gt;<br>
&gt;<br>
&gt; -- <br>
&gt; *Jacek Tomaka*<br>
&gt; Geophysical Software Developer<br>
&gt;<br>
&gt;<br>
&gt;<br>
&gt;<br>
&gt;<br>
&gt;<br>
&gt; *DownUnder GeoSolutions*<br>
&gt; 76 Kings Park Road<br>
&gt; West Perth 6005 WA, Australia<br>
&gt; *tel *+61 8 9287 4143 &lt;+61%208%209287%204143&gt;<br>
&gt; <a href="mailto:jacekt@dug.com" target="_blank">jacekt@dug.com</a><br>
&gt; *<a href="http://www.dug.com" rel="noreferrer" target="_blank">www.dug.com</a> &lt;<a href="http://www.dug.com" rel="noreferrer" target="_blank">http://www.dug.com</a>&gt;*<br>
</blockquote></div><br clear="all"><br>-- <br><div dir="ltr" class="gmail_signature"><div dir="ltr"><div><div dir="ltr"><div><div dir="ltr"><span><div><div dir="ltr"><div><div dir="ltr"><div><div dir="ltr"><div><font color="#000000"><font face="arial,helvetica,sans-serif"><b>Jacek Tomaka</b></font></font><br><font color="#000000"><font face="arial,helvetica,sans-serif"><font size="2">Geophysical Software Developer</font></font></font><br></div><font face="arial,helvetica,sans-serif" color="#000000">
</font><div><span lang="EN-US"></span> <span lang="EN-US"><b><br><br></b></span></div><font face="arial,helvetica,sans-serif" color="#000000">
</font><img src="http://drive.google.com/uc?export=view&amp;id=0B4X9ixpc-ZU_NHV0WnluaXp5ZkE"><br><br><span style="color:rgb(102,102,102)"><font size="2"><b>DownUnder GeoSolutions<br><br></b></font></span><div><span style="color:rgb(102,102,102)"></span><span style="color:rgb(102,102,102)">76 Kings Park Road<br></span></div><span style="color:rgb(102,102,102)">West Perth 6005 WA, Australia<br><i><b>tel </b></i><a href="tel:+61%208%209287%204143" value="+61892874143" target="_blank">+61 8 9287 4143</a><br><a href="mailto:jacekt@dug.com" target="_blank">jacekt@dug.com</a><br><b><a href="http://www.dug.com" target="_blank">www.dug.com</a></b></span></div></div></div></div></div></div></span></div></div></div></div></div></div></div>