<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=Windows-1252">
</head>
<body>
1. Good to know, thank you.&nbsp; I hadnt looked at the code, I was unaware it runs through all the sprinklers.<br>
<br>
2. Right, I know - the article was about it when it was an alias for reclaimable, and hence describes some of the behavior of reclaimable.<br>
<br>
3. Interesting, thats good to know.&nbsp; I would note that it doesnt seem to be standard practice in other file systems, though I didnt look at how many shrinkers theyre registering.&nbsp; Perhaps having special shrinkers is whats unusual.<br>
<br>
BTW, re: mailing list, this is the first devel appropriate thing Ive seen on discuss in a long while.&nbsp; I should instead have encouraged Jacek to use lustre-devel :)<br>
<br>
- Patrick
<hr style="display:inline-block;width:98%" tabindex="-1">
<div id="divRplyFwdMsg" dir="ltr"><font face="Calibri, sans-serif" style="font-size:11pt" color="#000000"><b>From:</b> NeilBrown &lt;neilb@suse.com&gt;<br>
<b>Sent:</b> Sunday, April 14, 2019 6:38:47 PM<br>
<b>To:</b> Patrick Farrell; Jacek Tomaka; lustre-discuss@lists.lustre.org<br>
<b>Subject:</b> Re: [lustre-discuss] Lustre client memory and MemoryAvailable</font>
<div>&nbsp;</div>
</div>
<div class="BodyFragment"><font size="2"><span style="font-size:11pt;">
<div class="PlainText"><br>
(that for the Cc Patrick - maybe I should subscribe to<br>
lustre-discuss...)<br>
<br>
1/ w.r.t drop_caches, &quot;2&quot; is *not* &quot;inode and dentry&quot;.&nbsp; The '2' bit<br>
&nbsp; causes all registered shrinkers to be run, until they report there is<br>
&nbsp; nothing left that can be discarded.&nbsp; If this is taking 10 minutes,<br>
&nbsp; then it seems likely that some shrinker is either very inefficient, or<br>
&nbsp; is reporting that there is more work to be done, when really there<br>
&nbsp; isn't.<br>
<br>
&nbsp; lustre registers 5 shrinkers.&nbsp; Any memcache which is not affected by<br>
&nbsp; those shrinkers should *not* be marked SLAB_RECLAIM_ACCOUNT (unless<br>
&nbsp; they are indirectly shrunk by a system shrinker - e.g. if they are<br>
&nbsp; slaves to the icache or dcache).&nbsp; Any which are probably can be.<br>
<br>
1a/ &quot;echo 3 &gt; drop_caches&quot; does the easy part of memory reclaim: it<br>
&nbsp;&nbsp; reclaims anything that can be reclaimed immediately.&nbsp; It doesn't<br>
&nbsp;&nbsp; trigger write-back, and it doesn't start the oom-killer, but all<br>
&nbsp;&nbsp; caches are flushed of everything that is not currently in use, and<br>
&nbsp;&nbsp; does not need to be written-out first.&nbsp; If you run &quot;sync&quot; first,<br>
&nbsp;&nbsp; there should be nothing to write out, so it should drop a lot more.<br>
<br>
2/ GFP_TEMPORARY is gone, it was never really well defined.&nbsp; Best to<br>
&nbsp;&nbsp; ignore it.<br>
<br>
3/ I don't *think* __GFP_RECLAIMABLE has a very big effect.&nbsp; It<br>
&nbsp;&nbsp; primarily tries to keep non-reclaimable allocations together so they<br>
&nbsp;&nbsp; don't cause too much fragmentation.&nbsp; To do this, it groups them<br>
&nbsp;&nbsp; separately from reclaimable allocations.&nbsp; So if a shrinker is<br>
&nbsp;&nbsp; expected to do anything useful, then it makes sense to tag the<br>
&nbsp;&nbsp; related slabs as RECLAIMABLE.<br>
<br>
4/ Patrick is right that accounting is best-effort.&nbsp; But we do want it<br>
&nbsp;&nbsp; to improve.&nbsp; Just last week there was a report<br>
&nbsp;&nbsp;&nbsp;&nbsp; <a href="https://lwn.net/SubscriberLink/784964/9ddad7d7050729e1/">https://lwn.net/SubscriberLink/784964/9ddad7d7050729e1/</a><br>
&nbsp;&nbsp; about making slab-allocated objects movable.&nbsp; If/when that gets off<br>
&nbsp;&nbsp; the ground, it should help the fragmentation problem, so more of the<br>
&nbsp;&nbsp; pages listed as reclaimable should actually be so.<br>
&nbsp;&nbsp; <br>
NeilBrown<br>
<br>
On Sun, Apr 14 2019, Patrick Farrell wrote:<br>
<br>
&gt; echo 1 &gt; drop_caches does not generate memory pressure - it requests that the page cache be cleared.&nbsp; It would not be expected to affect slab caches much.<br>
&gt;<br>
&gt; You could try 3 (1&#43;2 in this case, where 2 is inode and dentry).&nbsp; That might do a bit more because some (maybe many?) of those objects you're looking at would go away if the associated inodes or dentries were removed.&nbsp; But fundamentally, drop caches does
 not generate memory pressure, and does not force reclaim.&nbsp; It drops specific, identified caches.<br>
&gt;<br>
&gt; The only way to force *reclaim* is memory pressure.<br>
&gt;<br>
&gt; Your note that a lot more memory than expected was freed under pressure does tell us something, though.<br>
&gt;<br>
&gt; It's conceivable Lustre needs to set SLAB_RECLAIM_ACCOUNT on more of its slab caches, so this piqued my curiosity.&nbsp; My conclusion is no, here's why:<br>
&gt;<br>
&gt; The one quality reference I was quickly able to find suggests setting SLAB_RECLAIM_ACCOUNT wouldn't be so simple:<br>
&gt; <a href="https://lwn.net/Articles/713076/">https://lwn.net/Articles/713076/</a><br>
&gt;<br>
&gt; GFP_TEMPORARY is - in practice - just another name for __GFP_RECLAIMABLE, and setting SLAB_RECLAIM_ACCOUNT is equivalent to setting __GFP_RECLAIMABLE.&nbsp; That article suggests caution is needed, as this should only be used for memory that is certain to be easily
 available, because using this flag changes the allocation behavior on the assumption that the memory can be quickly freed at need.&nbsp; That is often not true of these Lustre objects.<br>
&gt;<br>
&gt; An easy way to learn more about this sort of question is to compare to other actively developed file systems in the kernel...<br>
&gt;<br>
&gt; Comparing to other file systems, we see that in general, only the inode cache is allocated with SLAB_RECLAIM_ACCOUNT (it varies a bit).<br>
&gt;<br>
&gt; XFS, for example, has only one use of KM_ZONE_RECLAIM, its name for this flag - the inode cache:<br>
&gt; &quot;<br>
&gt;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; xfs_inode_zone =<br>
&gt;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; kmem_zone_init_flags(sizeof(xfs_inode_t), &quot;xfs_inode&quot;,<br>
&gt;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; KM_ZONE_HWALIGN | KM_ZONE_RECLAIM | KM_ZONE_SPREAD,<br>
&gt;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; xfs_fs_inode_init_once);<br>
&gt; &quot;<br>
&gt;<br>
&gt; btrfs is the same, just the inode cache.&nbsp; EXT4 has a *few* more caches marked this way, but not everything.<br>
&gt;<br>
&gt; So, no - I don't think so.&nbsp; It would be atypical for Lustre to set SLAB_RECLAIM_ACCOUNT on its slab caches for internal objects.&nbsp; Presumably this sort of thing is not considered reclaimable enough for this purpose.<br>
&gt;<br>
&gt; I believe if you tried similar tests with other complex file systems (XFS might be a good start), you'd see broadly similar behavior.&nbsp; (Lustre is probably a bit worse because it has a more complex internal object model, so more slab caches.)<br>
&gt;<br>
&gt; VM accounting is distinctly imperfect.&nbsp; The design is such that it's often impossible to know how much memory could be made available without actually going and trying to free it.&nbsp; There are good, intrinsic reasons for some of that, and some of that is design
 artifacts...<br>
&gt;<br>
&gt; I've copied in Neil Brown, who I think only reads lustre-devel, just in case he has some particular input on this.<br>
&gt;<br>
&gt; Regards,<br>
&gt; - Patrick<br>
&gt; ________________________________<br>
&gt; From: lustre-discuss &lt;lustre-discuss-bounces@lists.lustre.org&gt; on behalf of Jacek Tomaka &lt;jacekt@dug.com&gt;<br>
&gt; Sent: Sunday, April 14, 2019 3:12:51 AM<br>
&gt; To: lustre-discuss@lists.lustre.org<br>
&gt; Subject: Re: [lustre-discuss] Lustre client memory and MemoryAvailable<br>
&gt;<br>
&gt; Actually i think it is just a bug with the way slab caches are created. Some of them should be passed a flag that they are reclaimable.<br>
&gt; i.e. something like:<br>
&gt; <a href="https://patchwork.kernel.org/patch/9360819/">https://patchwork.kernel.org/patch/9360819/</a><br>
&gt;<br>
&gt; Regards.<br>
&gt; Jacek Tomaka<br>
&gt;<br>
&gt; On Sun, Apr 14, 2019 at 3:27 PM Jacek Tomaka &lt;jacekt@dug.com&lt;mailto:jacekt@dug.com&gt;&gt; wrote:<br>
&gt; Hello,<br>
&gt;<br>
&gt; TL;DR;<br>
&gt; Is there a way to figure out how much memory Lustre will make available under memory pressure?<br>
&gt;<br>
&gt; Details:<br>
&gt; We are running lustre client on a machine with 128GB of memory (Centos 7) Intel Phi KNL machines and at certain situations we see that there can be ~10GB&#43; of memory allocated on the kernel side i.e. :<br>
&gt;<br>
&gt; vvp_object_kmem&nbsp;&nbsp; 3535336 3536986&nbsp;&nbsp;&nbsp; 176&nbsp;&nbsp; 46&nbsp;&nbsp;&nbsp; 2 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 : slabdata&nbsp; 76891&nbsp; 76891&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt; ll_thread_kmem&nbsp;&nbsp;&nbsp;&nbsp; 33511&nbsp; 33511&nbsp;&nbsp;&nbsp; 344&nbsp;&nbsp; 47&nbsp;&nbsp;&nbsp; 4 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 : slabdata&nbsp;&nbsp;&nbsp; 713&nbsp;&nbsp;&nbsp; 713&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt; lov_session_kmem&nbsp;&nbsp; 34760&nbsp; 34760&nbsp;&nbsp;&nbsp; 592&nbsp;&nbsp; 55&nbsp;&nbsp;&nbsp; 8 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 : slabdata&nbsp;&nbsp;&nbsp; 632&nbsp;&nbsp;&nbsp; 632&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt; osc_extent_kmem&nbsp;&nbsp; 3549831 3551232&nbsp;&nbsp;&nbsp; 168&nbsp;&nbsp; 48&nbsp;&nbsp;&nbsp; 2 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 : slabdata&nbsp; 73984&nbsp; 73984&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt; osc_thread_kmem&nbsp;&nbsp;&nbsp; 14012&nbsp; 14116&nbsp;&nbsp; 2832&nbsp;&nbsp; 11&nbsp;&nbsp;&nbsp; 8 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 : slabdata&nbsp;&nbsp; 1286&nbsp;&nbsp; 1286&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt; osc_object_kmem&nbsp;&nbsp; 3546640 3548350&nbsp;&nbsp;&nbsp; 304&nbsp;&nbsp; 53&nbsp;&nbsp;&nbsp; 4 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 : slabdata&nbsp; 66950&nbsp; 66950&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt; signal_cache&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 3702537 3707144&nbsp;&nbsp; 1152&nbsp;&nbsp; 28&nbsp;&nbsp;&nbsp; 8 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 : slabdata 132398 132398&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt;<br>
&gt; /proc/meminfo:<br>
&gt; MemAvailable:&nbsp;&nbsp; 114196044 kB<br>
&gt; Slab:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 11641808 kB<br>
&gt; SReclaimable:&nbsp;&nbsp;&nbsp; 1410732 kB<br>
&gt; SUnreclaim:&nbsp;&nbsp;&nbsp;&nbsp; 10231076 kB<br>
&gt;<br>
&gt; After executing<br>
&gt;<br>
&gt; echo 1 &gt;/proc/sys/vm/drop_caches<br>
&gt;<br>
&gt; the slabinfo values don't change but when i actually generate memory pressure by:<br>
&gt;<br>
&gt; java -Xmx117G -Xms117G -XX:&#43;AlwaysPreTouch -version<br>
&gt;<br>
&gt; lots of memory gets freed:<br>
&gt; vvp_object_kmem&nbsp;&nbsp; 127650 127880&nbsp;&nbsp;&nbsp; 176&nbsp;&nbsp; 46&nbsp;&nbsp;&nbsp; 2 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 : slabdata&nbsp;&nbsp; 2780&nbsp;&nbsp; 2780&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt; ll_thread_kmem&nbsp;&nbsp;&nbsp;&nbsp; 33558&nbsp; 33558&nbsp;&nbsp;&nbsp; 344&nbsp;&nbsp; 47&nbsp;&nbsp;&nbsp; 4 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 : slabdata&nbsp;&nbsp;&nbsp; 714&nbsp;&nbsp;&nbsp; 714&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt; lov_session_kmem&nbsp;&nbsp; 34815&nbsp; 34815&nbsp;&nbsp;&nbsp; 592&nbsp;&nbsp; 55&nbsp;&nbsp;&nbsp; 8 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 : slabdata&nbsp;&nbsp;&nbsp; 633&nbsp;&nbsp;&nbsp; 633&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt; osc_extent_kmem&nbsp;&nbsp; 128640 128880&nbsp;&nbsp;&nbsp; 168&nbsp;&nbsp; 48&nbsp;&nbsp;&nbsp; 2 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 : slabdata&nbsp;&nbsp; 2685&nbsp;&nbsp; 2685&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt; osc_thread_kmem&nbsp;&nbsp;&nbsp; 14038&nbsp; 14116&nbsp;&nbsp; 2832&nbsp;&nbsp; 11&nbsp;&nbsp;&nbsp; 8 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 : slabdata&nbsp;&nbsp; 1286&nbsp;&nbsp; 1286&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt; osc_object_kmem&nbsp;&nbsp;&nbsp; 82998&nbsp; 83263&nbsp;&nbsp;&nbsp; 304&nbsp;&nbsp; 53&nbsp;&nbsp;&nbsp; 4 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 : slabdata&nbsp;&nbsp; 1571&nbsp;&nbsp; 1571&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt; signal_cache&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 38734&nbsp; 44268&nbsp;&nbsp; 1152&nbsp;&nbsp; 28&nbsp;&nbsp;&nbsp; 8 : tunables&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp; 0 : slabdata&nbsp;&nbsp; 1581&nbsp;&nbsp; 1581&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<br>
&gt;<br>
&gt; /proc/meminfo:<br>
&gt; MemAvailable:&nbsp;&nbsp; 123146076 kB<br>
&gt; Slab:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 1959160 kB<br>
&gt; SReclaimable:&nbsp;&nbsp;&nbsp;&nbsp; 334276 kB<br>
&gt; SUnreclaim:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 1624884 kB<br>
&gt;<br>
&gt; The similar effect to generating memory pressure we see when executing:<br>
&gt;<br>
&gt; echo 3 &gt;/proc/sys/vm/drop_caches<br>
&gt;<br>
&gt; But this can take very long time (10 minutes).<br>
&gt;<br>
&gt; So essentially on a machine using Lustre client,&nbsp; MemAvailable is no longer a good predictor of the amount of memory that can be allocated.<br>
&gt; Is there a way to query Lustre and compensate for lustre cache memory that will be made available on memory pressure?<br>
&gt;<br>
&gt; Regards.<br>
&gt; --<br>
&gt; Jacek Tomaka<br>
&gt; Geophysical Software Developer<br>
&gt;<br>
&gt;<br>
&gt; [<a href="http://drive.google.com/uc?export=view&amp;id=0B4X9ixpc-ZU_NHV0WnluaXp5ZkE">http://drive.google.com/uc?export=view&amp;id=0B4X9ixpc-ZU_NHV0WnluaXp5ZkE</a>]<br>
&gt;<br>
&gt; DownUnder GeoSolutions<br>
&gt;<br>
&gt; 76 Kings Park Road<br>
&gt; West Perth 6005 WA, Australia<br>
&gt; tel &#43;61 8 9287 4143&lt;tel:&#43;61%208%209287%204143&gt;<br>
&gt; jacekt@dug.com&lt;<a href="mailto:jacekt@dug.com">mailto:jacekt@dug.com</a>&gt;<br>
&gt; www.dug.com&lt;<a href="http://www.dug.com">http://www.dug.com</a>&gt;<br>
&gt;<br>
&gt;<br>
&gt; --<br>
&gt; Jacek Tomaka<br>
&gt; Geophysical Software Developer<br>
&gt;<br>
&gt;<br>
&gt; [<a href="http://drive.google.com/uc?export=view&amp;id=0B4X9ixpc-ZU_NHV0WnluaXp5ZkE">http://drive.google.com/uc?export=view&amp;id=0B4X9ixpc-ZU_NHV0WnluaXp5ZkE</a>]<br>
&gt;<br>
&gt; DownUnder GeoSolutions<br>
&gt;<br>
&gt; 76 Kings Park Road<br>
&gt; West Perth 6005 WA, Australia<br>
&gt; tel &#43;61 8 9287 4143&lt;tel:&#43;61%208%209287%204143&gt;<br>
&gt; jacekt@dug.com&lt;<a href="mailto:jacekt@dug.com">mailto:jacekt@dug.com</a>&gt;<br>
&gt; www.dug.com&lt;<a href="http://www.dug.com">http://www.dug.com</a>&gt;<br>
</div>
</span></font></div>
</body>
</html>