Is this filesystem nearly full? Fragmentation can decrease back end performance.<br><br>Also check the disks stats on the DDN, maybe you have a slow disk in one of your tiers.<br><br>Wojciech<br><br><div class="gmail_quote">
On 18 October 2010 18:49, Peter Kjellstrom <span dir="ltr">&lt;<a href="mailto:cap@nsc.liu.se">cap@nsc.liu.se</a>&gt;</span> wrote:<br><blockquote class="gmail_quote" style="margin: 0pt 0pt 0pt 0.8ex; border-left: 1px solid rgb(204, 204, 204); padding-left: 1ex;">
<div class="im">On Monday 18 October 2010, John White wrote:<br>
&gt; Hello Folks,<br>
&gt;    A while back (say 3 weeks ago) we started noticing extremely high loads<br>
&gt; (load avg around 300 at times) on our OSSs when in production and serving<br>
&gt; IO. This cluster was, at the time, on 1.8.2 (we have since upgraded to<br>
&gt; 1.8.4 but the problem remains). The load increases fairly predictably as<br>
&gt; clients generate IO but even 2 clients can produce a load avg above 5.00.<br>
<br>
</div>Does this impact performance or does it only show up as an unexpectedly high<br>
number on the OSSes?<br>
<font color="#888888"><br>
/Peter<br>
</font><div><div></div><div class="h5"><br>
&gt; An identical file system of ours does not exhibit this behavior (sticks<br>
&gt; below load avg 1.00 under even the heaviest IO load). I&#39;ve looked around<br>
&gt; bugzilla and haven&#39;t found anything. We&#39;ve disabled heartbeat on the<br>
&gt; off-chance that was generating the load (it&#39;s not), we&#39;ve attempted using a<br>
&gt; different client transport (o2ib-&gt;tcp), this did not solve the issue.<br>
&gt; There doesn&#39;t appear to be any specific non-kernel thread causing the<br>
&gt; high-load. The only info in dmesg/syslog pertains to sporadic client<br>
&gt; evictions or sporadic slow setattr due to heavy IO load (we&#39;ve since tuned<br>
&gt; the number of OST threads). We&#39;re basically out of ideas to try.<br>
&gt;<br>
&gt; As reference, this is a 1 MDS/4 OSS cluster backed by a DDN 9900 couplet<br>
&gt; (15 tiers, 1:1 lun mapping) running the <a href="http://lustre.org" target="_blank">lustre.org</a> rpm build kernel for<br>
&gt; 1.8.4. The MDS/OSSs are Dell R710s and the MDT is a Dell MD1000. Is this<br>
&gt; a common problem or should a bug be filed? Any info available upon<br>
&gt; request. Thanks for your time. ----------------<br>
&gt; John White<br>
&gt; High Performance Computing Services (HPCS)<br>
&gt; (510) 486-7307<br>
&gt; One Cyclotron Rd, MS: 50B-3209C<br>
&gt; Lawrence Berkeley National Lab<br>
&gt; Berkeley, CA 94720<br>
</div></div><br>_______________________________________________<br>
Lustre-discuss mailing list<br>
<a href="mailto:Lustre-discuss@lists.lustre.org">Lustre-discuss@lists.lustre.org</a><br>
<a href="http://lists.lustre.org/mailman/listinfo/lustre-discuss" target="_blank">http://lists.lustre.org/mailman/listinfo/lustre-discuss</a><br>
<br></blockquote></div><br><br clear="all"><br>-- <br>Wojciech Turek<br><br>Senior System Architect<br><br>High Performance Computing Service<br>University of Cambridge<br>Email: <a href="mailto:wjt27@cam.ac.uk" target="_blank">wjt27@cam.ac.uk</a><br>
Tel: (+)44 1223 763517 <br>