<div dir="ltr"><br><div class="gmail_extra"><br><div class="gmail_quote">On Wed, Sep 2, 2015 at 8:47 PM, Wahl, Edward <span dir="ltr">&lt;<a href="mailto:ewahl@osc.edu" target="_blank">ewahl@osc.edu</a>&gt;</span> wrote:<br><blockquote class="gmail_quote" style="margin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex">




<div>
<div style="direction:ltr;font-family:Tahoma;color:#000000;font-size:10pt">I&#39;ve seen this kind of error before when doing samba to do something stupid (and let&#39;s face it, that most everything with samba)   It was a locking issue I think.   Things were
 being changed/deleted/ (unlinked in actuality)  as the client was trying to do something with it. 
<br>
<br>
 Is the Apache process or it&#39;s spawned app(s)  still working on the files in question while serving them up?<br></div></div></blockquote><div>Not as far as I know, these are result files that were generated days ago (possibly more) and should be static by now....<br></div><div>But I&#39;ll double check with the people behind the app....  <br></div><blockquote class="gmail_quote" style="margin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex"><div><div style="direction:ltr;font-family:Tahoma;color:#000000;font-size:10pt">
That would be my guess here.  Any chance this is across NFS?  Seen that a great deal with this error, it used to cause crashes.<br></div></div></blockquote><div>Strictly speaking it is not, but it may be because a part of the path the server &#39;sees&#39;/&#39;knows&#39; is a symlink to the lustre filesystem which lives on nfs...<br><br><br></div><div>Thanks,<br></div><div>Eli<br></div><blockquote class="gmail_quote" style="margin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex"><div><div style="direction:ltr;font-family:Tahoma;color:#000000;font-size:10pt">
<br>
Ed Wahl<br>
OSC<br>
<br>
<br>
<div style="font-family:Times New Roman;color:#000000;font-size:16px">
<hr>
<div style="direction:ltr"><font face="Tahoma" size="2" color="#000000"><b>From:</b> lustre-discuss [<a href="mailto:lustre-discuss-bounces@lists.lustre.org" target="_blank">lustre-discuss-bounces@lists.lustre.org</a>] on behalf of E.S. Rosenberg [<a href="mailto:esr%2Blustre@mail.hebrew.edu" target="_blank">esr+lustre@mail.hebrew.edu</a>]<br>
<b>Sent:</b> Wednesday, September 02, 2015 7:57 AM<br>
<b>To:</b> <a href="mailto:lustre-discuss@lists.lustre.org" target="_blank">lustre-discuss@lists.lustre.org</a><br>
<b>Subject:</b> [lustre-discuss] refresh file layout error<br>
</font><br>
</div><div><div class="h5">
<div></div>
<div>
<div dir="ltr">
<div>
<div>Hi all,<br>
<br>
</div>
I am seeing an interesting/annoying problem with lustre and am not really sure what/where to look.<br>
<br>
</div>
When a webserver (galaxy using wsgi/apache2) tries to server (large) files stored on lustre it fails to send the full file and I see the following errors in syslog:<br>
<div>
<div>
<div><br>
Sep  2 11:50:17 hm-02 kernel: LustreError: 6973:0:(vvp_io.c:1197:vvp_io_init()) fs01: refresh file layout [0x200008815:0x217e:0x0] error -13.<br>
Sep  2 11:50:17 hm-02 kernel: LustreError: 6973:0:(file.c:179:ll_close_inode_openhandle()) inode 144115772543738238 mdc close failed: rc = -13<br>
<br>
</div>
<div>If I try to access the files through their direct path (copying to tmp/md5sum/sha512sum) it seems to work without a problem (full file is copied and sums agree, from different nodes).<br>
<br>
</div>
<div>When we switched the storage backend to NFS the server worked fine, so my guess is that there is an issue with the way python tries to read from the &#39;disk&#39;...<br>
<br>
</div>
<div>Is anyone familiar with the error above?<br>
<br>
</div>
<div>Thanks,<br>
</div>
<div>Eli<br>
</div>
</div>
</div>
</div>
</div>
</div></div></div>
</div>
</div>

</blockquote></div><br></div></div>