<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=us-ascii">
</head>
<body>
Neil,<br>
<br>
My understanding is marking the inode cache reclaimable would make Lustre unusual/unique among Linux file systems.&nbsp; Is that incorrect?<br>
<br>
- Patrick
<hr style="display:inline-block;width:98%" tabindex="-1">
<div id="divRplyFwdMsg" dir="ltr"><font face="Calibri, sans-serif" style="font-size:11pt" color="#000000"><b>From:</b> lustre-discuss &lt;lustre-discuss-bounces@lists.lustre.org&gt; on behalf of NeilBrown &lt;neilb@suse.com&gt;<br>
<b>Sent:</b> Monday, April 29, 2019 8:53:43 PM<br>
<b>To:</b> Jacek Tomaka<br>
<b>Cc:</b> lustre-discuss@lists.lustre.org<br>
<b>Subject:</b> Re: [lustre-discuss] Lustre client memory and MemoryAvailable</font>
<div>&nbsp;</div>
</div>
<div class="BodyFragment"><font size="2"><span style="font-size:11pt;">
<div class="PlainText">On Mon, Apr 29 2019, Jacek Tomaka wrote:<br>
<br>
&gt;&gt; so lustre_inode_cache is the real culprit when signal_cache appears to<br>
&gt;&gt;&nbsp; be large.<br>
&gt;&gt; This cache is slaved on the common inode cache, so there should be one<br>
&gt;&gt; entry for each lustre inode that is in memory.<br>
&gt;&gt; These inodes should get pruned when they've been inactive for a while.<br>
&gt;<br>
&gt; What triggers the prunning?<br>
&gt;<br>
<br>
Memory pressure.<br>
The approx approach is try to free some unused pages and about 1/2000th of<br>
the entries in each slab.&nbsp; Then if that hasn't made enough space<br>
available, try again.<br>
<br>
&gt;&gt;If you look in /proc/sys/fs/inode-nr&nbsp; there should be two numbers:<br>
&gt;&gt;&nbsp; The first is the total number of in-memory inodes for all filesystems.<br>
&gt;&gt;&nbsp; The second is the number of &quot;unused&quot; inodes.<br>
&gt;&gt;<br>
&gt;&gt;&nbsp; When you write &quot;3&quot; to drop_caches, the second number should drop down to<br>
&gt;&gt; nearly zero (I get 95 on my desktop, down from 6524).<br>
&gt;<br>
&gt; Ok, that is useful to know but echoing 3 to drop_cache or generating memory<br>
&gt; pressure<br>
&gt; clears most of the signal_cache (inode) as well as other lustre objects, so<br>
&gt; this is working fine.<br>
<br>
Oh good, I hadn't remembered clearly what the issue was.<br>
<br>
&gt;<br>
&gt; The issue that remains is that they are marked as SUnreclaim vs<br>
&gt; SReclaimable.<br>
<br>
Yes, I think lustre_inode_cache should certainly be flagged as<br>
SLAB_RECLAIM_ACCOUNT.<br>
If the SReclaimable value is too small (and there aren't many<br>
reclaimable pagecache pages), vmscan can decide not to bother.&nbsp; This is<br>
probably a fairly small risk but it is possible that the missing<br>
SLAB_RECLAIM_ACCOUNT flag can result in memory not being reclaimed when<br>
it could be.<br>
<br>
Thanks,<br>
NeilBrown<br>
<br>
<br>
&gt; So i do not think there is a memory leak per se.<br>
&gt;<br>
&gt; Regards.<br>
&gt; Jacek Tomaka<br>
&gt;<br>
&gt; On Mon, Apr 29, 2019 at 1:39 PM NeilBrown &lt;neilb@suse.com&gt; wrote:<br>
&gt;<br>
&gt;&gt;<br>
&gt;&gt; Thanks Jacek,<br>
&gt;&gt;&nbsp; so lustre_inode_cache is the real culprit when signal_cache appears to<br>
&gt;&gt;&nbsp; be large.<br>
&gt;&gt;&nbsp; This cache is slaved on the common inode cache, so there should be one<br>
&gt;&gt;&nbsp; entry for each lustre inode that is in memory.<br>
&gt;&gt;&nbsp; These inodes should get pruned when they've been inactive for a while.<br>
&gt;&gt;<br>
&gt;&gt;&nbsp; If you look in /proc/sys/fs/inode-nr&nbsp; there should be two numbers:<br>
&gt;&gt;&nbsp;&nbsp; The first is the total number of in-memory inodes for all filesystems.<br>
&gt;&gt;&nbsp;&nbsp; The second is the number of &quot;unused&quot; inodes.<br>
&gt;&gt;<br>
&gt;&gt;&nbsp; When you write &quot;3&quot; to drop_caches, the second number should drop down to<br>
&gt;&gt;&nbsp; nearly zero (I get 95 on my desktop, down from 6524).<br>
&gt;&gt;<br>
&gt;&gt;&nbsp; When signal_cache stays large even after the drop_caches, it suggest<br>
&gt;&gt;&nbsp; that there are lots of lustre inodes that are thought to be still<br>
&gt;&gt;&nbsp; active.&nbsp;&nbsp; I'd have to do a bit of digging to understand what that means,<br>
&gt;&gt;&nbsp; and a lot more to work out why lustre is holding on to inodes longer<br>
&gt;&gt;&nbsp; than you would expect (if that actually is the case).<br>
&gt;&gt;<br>
&gt;&gt;&nbsp; If an inode still has cached data pages attached that cannot easily be<br>
&gt;&gt;&nbsp; removed, it will not be purged even if it is unused.<br>
&gt;&gt;&nbsp; So if you see the &quot;unused&quot; number remaining high even after a<br>
&gt;&gt;&nbsp; &quot;drop_caches&quot;, that might mean that lustre isn't letting go of cache<br>
&gt;&gt;&nbsp; pages for some reason.<br>
&gt;&gt;<br>
&gt;&gt; NeilBrown<br>
&gt;&gt;<br>
&gt;&gt;<br>
&gt;&gt;<br>
&gt;&gt; On Mon, Apr 29 2019, Jacek Tomaka wrote:<br>
&gt;&gt;<br>
&gt;&gt; &gt; Wow, Thanks Nathan and NeilBrown.<br>
&gt;&gt; &gt; It is great to learn about slub merging. It is awesome to have a<br>
&gt;&gt; &gt; reproducer.<br>
&gt;&gt; &gt; I am yet to trigger my original problem with slurm_nomerge but<br>
&gt;&gt; &gt; slabinfo tool (in kernel sources) can actually show merged caches:<br>
&gt;&gt; &gt; kernel/3.10.0-693.5.2.el7/tools/slabinfo&nbsp; -a<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; :t-0000112&nbsp;&nbsp; &lt;- sysfs_dir_cache kernfs_node_cache blkdev_integrity<br>
&gt;&gt; &gt; task_delay_info<br>
&gt;&gt; &gt; :t-0000144&nbsp;&nbsp; &lt;- flow_cache cl_env_kmem<br>
&gt;&gt; &gt; :t-0000160&nbsp;&nbsp; &lt;- sigqueue lov_object_kmem<br>
&gt;&gt; &gt; :t-0000168&nbsp;&nbsp; &lt;- lovsub_object_kmem osc_extent_kmem<br>
&gt;&gt; &gt; :t-0000176&nbsp;&nbsp; &lt;- vvp_object_kmem nfsd4_stateids<br>
&gt;&gt; &gt; :t-0000192&nbsp;&nbsp; &lt;- ldlm_resources kiocb cred_jar inet_peer_cache key_jar<br>
&gt;&gt; &gt; file_lock_cache kmalloc-192 dmaengine-unmap-16 bio_integrity_payload<br>
&gt;&gt; &gt; :t-0000216&nbsp;&nbsp; &lt;- vvp_session_kmem vm_area_struct<br>
&gt;&gt; &gt; :t-0000256&nbsp;&nbsp; &lt;- biovec-16 ip_dst_cache bio-0 ll_file_data kmalloc-256<br>
&gt;&gt; &gt; sgpool-8 filp request_sock_TCP rpc_tasks request_sock_TCPv6<br>
&gt;&gt; &gt; skbuff_head_cache pool_workqueue lov_thread_kmem<br>
&gt;&gt; &gt; :t-0000264&nbsp;&nbsp; &lt;- osc_lock_kmem numa_policy<br>
&gt;&gt; &gt; :t-0000328&nbsp;&nbsp; &lt;- osc_session_kmem taskstats<br>
&gt;&gt; &gt; :t-0000576&nbsp;&nbsp; &lt;- kioctx xfrm_dst_cache vvp_thread_kmem<br>
&gt;&gt; &gt; :t-0001152&nbsp;&nbsp; &lt;- signal_cache lustre_inode_cache<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; It is not on a machine that had the problem i described before but the<br>
&gt;&gt; &gt; kernel version is the same so I am assuming the cache merges are the<br>
&gt;&gt; same.<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; Looks like signal_cache points to lustre_inode_cache.<br>
&gt;&gt; &gt; Regards.<br>
&gt;&gt; &gt; Jacek Tomaka<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; On Thu, Apr 25, 2019 at 7:42 AM NeilBrown &lt;neilb@suse.com&gt; wrote:<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt;<br>
&gt;&gt; &gt;&gt; Hi,<br>
&gt;&gt; &gt;&gt;&nbsp; you seem to be able to reproduce this fairly easily.<br>
&gt;&gt; &gt;&gt;&nbsp; If so, could you please boot with the &quot;slub_nomerge&quot; kernel parameter<br>
&gt;&gt; &gt;&gt;&nbsp; and then reproduce the (apparent) memory leak.<br>
&gt;&gt; &gt;&gt;&nbsp; I'm hoping that this will show some other slab that is actually using<br>
&gt;&gt; &gt;&gt;&nbsp; the memory - a slab with very similar object-size to signal_cache that<br>
&gt;&gt; &gt;&gt;&nbsp; is, by default, being merged with signal_cache.<br>
&gt;&gt; &gt;&gt;<br>
&gt;&gt; &gt;&gt; Thanks,<br>
&gt;&gt; &gt;&gt; NeilBrown<br>
&gt;&gt; &gt;&gt;<br>
&gt;&gt; &gt;&gt;<br>
&gt;&gt; &gt;&gt; On Wed, Apr 24 2019, Nathan Dauchy - NOAA Affiliate wrote:<br>
&gt;&gt; &gt;&gt;<br>
&gt;&gt; &gt;&gt; &gt; On Mon, Apr 15, 2019 at 9:18 PM Jacek Tomaka &lt;jacekt@dug.com&gt; wrote:<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt;&gt;<br>
&gt;&gt; &gt;&gt; &gt;&gt; &gt;signal_cache should have one entry for each process (or<br>
&gt;&gt; thread-group).<br>
&gt;&gt; &gt;&gt; &gt;&gt;<br>
&gt;&gt; &gt;&gt; &gt;&gt; That is what i thought as well, looking at the kernel source,<br>
&gt;&gt; &gt;&gt; allocations<br>
&gt;&gt; &gt;&gt; &gt;&gt; from<br>
&gt;&gt; &gt;&gt; &gt;&gt; signal_cache happen only during fork.<br>
&gt;&gt; &gt;&gt; &gt;&gt;<br>
&gt;&gt; &gt;&gt; &gt;&gt;<br>
&gt;&gt; &gt;&gt; &gt; I was recently chasing an issue with clients suffering from low memory<br>
&gt;&gt; &gt;&gt; and<br>
&gt;&gt; &gt;&gt; &gt; saw that &quot;signal_cache&quot; was a major player.&nbsp; But the workload on those<br>
&gt;&gt; &gt;&gt; &gt; clients was not doing a lot of forking.&nbsp; (and I don't *think*<br>
&gt;&gt; threading<br>
&gt;&gt; &gt;&gt; &gt; either)&nbsp; Rather it was a LOT of metadata read operations.<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; You can see the symptoms by a simple &quot;du&quot; on a Lustre file system:<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; # grep signal_cache /proc/slabinfo<br>
&gt;&gt; &gt;&gt; &gt; signal_cache&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 967&nbsp;&nbsp; 1092&nbsp;&nbsp; 1152&nbsp;&nbsp; 28&nbsp;&nbsp;&nbsp; 8 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0<br>
&gt;&gt; &gt;&gt; 0<br>
&gt;&gt; &gt;&gt; &gt; : slabdata&nbsp;&nbsp;&nbsp;&nbsp; 39&nbsp;&nbsp;&nbsp;&nbsp; 39&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; # du -s /mnt/lfs1/projects/foo<br>
&gt;&gt; &gt;&gt; &gt; 339744908 /mnt/lfs1/projects/foo<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; # grep signal_cache /proc/slabinfo<br>
&gt;&gt; &gt;&gt; &gt; signal_cache&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 164724 164724&nbsp;&nbsp; 1152&nbsp;&nbsp; 28&nbsp;&nbsp;&nbsp; 8 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0<br>
&gt;&gt; &gt;&gt; 0<br>
&gt;&gt; &gt;&gt; &gt; : slabdata&nbsp;&nbsp; 5883&nbsp;&nbsp; 5883&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; # slabtop -s c -o | head -n 20<br>
&gt;&gt; &gt;&gt; &gt;&nbsp; Active / Total Objects (% used)&nbsp;&nbsp;&nbsp; : 3660791 / 3662863 (99.9%)<br>
&gt;&gt; &gt;&gt; &gt;&nbsp; Active / Total Slabs (% used)&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; : 93019 / 93019 (100.0%)<br>
&gt;&gt; &gt;&gt; &gt;&nbsp; Active / Total Caches (% used)&nbsp;&nbsp;&nbsp;&nbsp; : 72 / 107 (67.3%)<br>
&gt;&gt; &gt;&gt; &gt;&nbsp; Active / Total Size (% used)&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; : 836474.91K / 837502.16K (99.9%)<br>
&gt;&gt; &gt;&gt; &gt;&nbsp; Minimum / Average / Maximum Object : 0.01K / 0.23K / 12.75K<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt;&nbsp;&nbsp; OBJS ACTIVE&nbsp; USE OBJ SIZE&nbsp; SLABS OBJ/SLAB CACHE SIZE NAME<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; 164724 164724 100%&nbsp;&nbsp;&nbsp; 1.12K&nbsp;&nbsp; 5883&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 28&nbsp;&nbsp;&nbsp; 188256K signal_cache<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; 331712 331712 100%&nbsp;&nbsp;&nbsp; 0.50K&nbsp; 10366&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 32&nbsp;&nbsp;&nbsp; 165856K ldlm_locks<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; 656896 656896 100%&nbsp;&nbsp;&nbsp; 0.12K&nbsp; 20528&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 32&nbsp;&nbsp;&nbsp;&nbsp; 82112K kmalloc-128<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; 340200 339971&nbsp; 99%&nbsp;&nbsp;&nbsp; 0.19K&nbsp;&nbsp; 8100&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 42&nbsp;&nbsp;&nbsp;&nbsp; 64800K kmalloc-192<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; 162838 162838 100%&nbsp;&nbsp;&nbsp; 0.30K&nbsp;&nbsp; 6263&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 26&nbsp;&nbsp;&nbsp;&nbsp; 50104K osc_object_kmem<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; 744192 744192 100%&nbsp;&nbsp;&nbsp; 0.06K&nbsp; 11628&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 64&nbsp;&nbsp;&nbsp;&nbsp; 46512K kmalloc-64<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; 205128 205128 100%&nbsp;&nbsp;&nbsp; 0.19K&nbsp;&nbsp; 4884&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 42&nbsp;&nbsp;&nbsp;&nbsp; 39072K dentry<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt;&nbsp;&nbsp; 4268&nbsp;&nbsp; 4256&nbsp; 99%&nbsp;&nbsp;&nbsp; 8.00K&nbsp;&nbsp; 1067&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 4&nbsp;&nbsp;&nbsp;&nbsp; 34144K kmalloc-8192<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; 162978 162978 100%&nbsp;&nbsp;&nbsp; 0.17K&nbsp;&nbsp; 3543&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 46&nbsp;&nbsp;&nbsp;&nbsp; 28344K vvp_object_kmem<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; 162792 162792 100%&nbsp;&nbsp;&nbsp; 0.16K&nbsp;&nbsp; 6783&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 24&nbsp;&nbsp;&nbsp;&nbsp; 27132K<br>
&gt;&gt; &gt;&gt; kvm_mmu_page_header<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; 162825 162825 100%&nbsp;&nbsp;&nbsp; 0.16K&nbsp;&nbsp; 6513&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 25&nbsp;&nbsp;&nbsp;&nbsp; 26052K sigqueue<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt;&nbsp; 16368&nbsp; 16368 100%&nbsp;&nbsp;&nbsp; 1.02K&nbsp;&nbsp;&nbsp; 528&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 31&nbsp;&nbsp;&nbsp;&nbsp; 16896K nfs_inode_cache<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt;&nbsp; 20385&nbsp; 20385 100%&nbsp;&nbsp;&nbsp; 0.58K&nbsp;&nbsp;&nbsp; 755&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 27&nbsp;&nbsp;&nbsp;&nbsp; 12080K inode_cache<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; Repeat that for more (and bigger) directories and slab cache added up<br>
&gt;&gt; to<br>
&gt;&gt; &gt;&gt; &gt; more than half the memory on this 24GB node.<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; This is with CentOS-7.6 and lustre-2.10.5_ddn6.<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; I worked around the problem by tackling the &quot;ldlm_locks&quot; memory usage<br>
&gt;&gt; &gt;&gt; with:<br>
&gt;&gt; &gt;&gt; &gt; # lctl set_param ldlm.namespaces.lfs*.lru_max_age=10000<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; ...but I did not find a way to reduce the &quot;signal_cache&quot;.<br>
&gt;&gt; &gt;&gt; &gt;<br>
&gt;&gt; &gt;&gt; &gt; Regards,<br>
&gt;&gt; &gt;&gt; &gt; Nathan<br>
&gt;&gt; &gt;&gt;<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; --<br>
&gt;&gt; &gt; *Jacek Tomaka*<br>
&gt;&gt; &gt; Geophysical Software Developer<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt;<br>
&gt;&gt; &gt; *DownUnder GeoSolutions*<br>
&gt;&gt; &gt; 76 Kings Park Road<br>
&gt;&gt; &gt; West Perth 6005 WA, Australia<br>
&gt;&gt; &gt; *tel *&#43;61 8 9287 4143 &lt;&#43;61%208%209287%204143&gt;<br>
&gt;&gt; &gt; jacekt@dug.com<br>
&gt;&gt; &gt; *www.dug.com &lt;<a href="http://www.dug.com">http://www.dug.com</a>&gt;*<br>
&gt;&gt;<br>
&gt;<br>
&gt;<br>
&gt; -- <br>
&gt; *Jacek Tomaka*<br>
&gt; Geophysical Software Developer<br>
&gt;<br>
&gt;<br>
&gt;<br>
&gt;<br>
&gt;<br>
&gt;<br>
&gt; *DownUnder GeoSolutions*<br>
&gt; 76 Kings Park Road<br>
&gt; West Perth 6005 WA, Australia<br>
&gt; *tel *&#43;61 8 9287 4143 &lt;&#43;61%208%209287%204143&gt;<br>
&gt; jacekt@dug.com<br>
&gt; *www.dug.com &lt;<a href="http://www.dug.com">http://www.dug.com</a>&gt;*<br>
</div>
</span></font></div>
</body>
</html>