Hi Bernd,<br><br>Many thanks for your reply. I have found this bug last night and as far as I can see there is no fix for it yet? I am preparing dbs to run lfsck on affected file systems. I also found bug 18748 and I must say we have exactly the same problems. It just looks like we run into that problem few months after CIEMAT did. As far as I know if we can see this message it means that there are files with missing objects. The worst is that we don&#39;t know when and why files looses they objects. It just happens spontaneously and there isn&#39;t any lustre messages that could give us a clue. Users run jobs and some time after their files were written some of these files get corrupted/looses objects (?-----) trying to access this files for the first time triggers &#39;lvbo&#39; message.<br>
We have third lustre file system which runs on different hardware but the same lustre version and RHEL version as the affected ones. I can not see any problems on the third file system.<br><br>Wojciech  <br><br><div class="gmail_quote">
2009/10/10 Bernd Schubert <span dir="ltr">&lt;<a href="mailto:bs_lists@aakef.fastmail.fm" target="_blank">bs_lists@aakef.fastmail.fm</a>&gt;</span><br><blockquote class="gmail_quote" style="border-left: 1px solid rgb(204, 204, 204); margin: 0pt 0pt 0pt 0.8ex; padding-left: 1ex;">

&quot;ASSERTION(old_inode-&gt;i_state &amp; I_FREEING)&quot; is the infamous bug17485. You will<br>
need to run lfsck to fix it.<br>
<div><div></div><div><br>
<br>
On Saturday 10 October 2009, Wojciech Turek wrote:<br>
&gt; Hi,<br>
&gt;<br>
&gt; Did you get to the bottom of this?<br>
&gt;<br>
&gt; We are having exactly the same problem with our lustre-1.6.6 (rhel4) file<br>
&gt; systems. Recently it got worst and MDS crashes quite frequently, when we<br>
&gt; run e2fsck there are errors that are being fixed. However after some time<br>
&gt; we still are seeing the same errors in the logs about missing objects and<br>
&gt; files get corrupted (?-----------) Also clients LBUGs quite frequently<br>
&gt; with this message (osc_request.c:2904:osc_set_data_with_check()) LBUG<br>
&gt; This looks like serious lustre problem but so far I didn&#39;t find any clues<br>
&gt; on that even after long search through lustre bugzilla.<br>
&gt;<br>
&gt; Our MDSs and OSSs are UPSed, RAID is behaving OK, we don&#39;t see any errors<br>
&gt; in the syslog.<br>
&gt;<br>
&gt; I will be grateful for some hints on this one<br>
&gt;<br>
&gt; Wojciech<br>
&gt;<br>
&gt; 2009/8/24 rishi pathak &lt;<a href="mailto:mailmaverick666@gmail.com" target="_blank">mailmaverick666@gmail.com</a>&gt;<br>
&gt;<br>
&gt; &gt; Hi,<br>
&gt; &gt;<br>
&gt; &gt; Our lustre fs comprises of 15 OST/OSS and 1 MDS with no failover. Client<br>
&gt; &gt; as well as servers run lustre-1.6 and kernel 2.6.9-18.<br>
&gt; &gt;<br>
&gt; &gt;    Doing a ls -ltr for a directory in lustre fs throws following<br>
&gt; &gt; errors (as got from lustre logs) on client<br>
&gt; &gt;<br>
&gt; &gt; 00000008:00020000:0:1251099455.304622:0:724:0:(osc_request.c:2898:osc_set<br>
&gt; &gt;_data_with_check()) ### inconsistent l_ast_data found ns:<br>
&gt; &gt; scratch-OST0005-osc-ffff81201e8dd800 lock: ffff811f9af04<br>
&gt; &gt; 000/0xec0d1c36da6992fd lrc: 3/1,0 mode: PR/PR res: 570622/0 rrc: 2 type:<br>
&gt; &gt; EXT [0-&gt;18446744073709551615] (req 0-&gt;18446744073709551615) flags: 100000<br>
&gt; &gt; remote: 0xb79b445e381bc9e6 expref: -99 p<br>
&gt; &gt; id: 22878<br>
&gt; &gt; 00000008:00040000:0:1251099455.337868:0:724:0:(osc_request.c:2904:osc_set<br>
&gt; &gt;_data_with_check()) ASSERTION(old_inode-&gt;i_state &amp; I_FREEING) failed:Found<br>
&gt; &gt; existing inode ffff811f2cf693b8/1972725<br>
&gt; &gt; 44/1895600178 state 0 in lock: setting data to<br>
&gt; &gt; ffff8118ef8ed5f8/207519777/1771835328<br>
&gt; &gt; 00000000:00040000:0:1251099455.360090:0:724:0:(osc_request.c:2904:osc_set<br>
&gt; &gt;_data_with_check()) LBUG<br>
&gt; &gt;<br>
&gt; &gt;<br>
&gt; &gt; On scratch-OST0005 OST it shows<br>
&gt; &gt;<br>
&gt; &gt; Aug 24 10:22:53 yn266 kernel: LustreError:<br>
&gt; &gt; 3023:0:(ldlm_resource.c:851:ldlm_resource_add()) lvbo_init failed for<br>
&gt; &gt; resour ce 569204: rc -2<br>
&gt; &gt; Aug 24 10:22:53 yn266 kernel: LustreError:<br>
&gt; &gt; 3023:0:(ldlm_resource.c:851:ldlm_resource_add()) Skipped 19 previous<br>
&gt; &gt; similar messages<br>
&gt; &gt; Aug 24 12:40:43 yn266 kernel: LustreError:<br>
&gt; &gt; 2737:0:(ldlm_resource.c:851:ldlm_resource_add()) lvbo_init failed for<br>
&gt; &gt; resour ce 569195: rc -2<br>
&gt; &gt; Aug 24 12:44:59 yn266 kernel: LustreError:<br>
&gt; &gt; 2835:0:(ldlm_resource.c:851:ldlm_resource_add()) lvbo_init failed for<br>
&gt; &gt; resour ce 569198: rc -2<br>
&gt; &gt;<br>
&gt; &gt; These kind of errors we are getting for many clients.<br>
&gt; &gt;<br>
&gt; &gt; ##History ##<br>
&gt; &gt; Prior to thsese occurences, our MDS showed signs of failure in way that<br>
&gt; &gt; cpu load was shooting above 100 (on a quad core quad socket system) and<br>
&gt; &gt; users were complaining about slow storage performance. We took it offline<br>
&gt; &gt; and did fsck on unmounted MDS and OSTs. fsck on OSTs went fine but it<br>
&gt; &gt; showed some errors which were fixed. For data integrity check, mdsdb and<br>
&gt; &gt; ostdb were built and lfsck was run on a client(client was mounted with<br>
&gt; &gt; abort_recov).<br>
&gt; &gt;<br>
&gt; &gt; lfsck was run in following order:<br>
&gt; &gt; lfsck with no fix - reported dangling inodes and orphaned objects<br>
&gt; &gt; lfsck with -l (backup orphaned objects)<br>
&gt; &gt; lfsck with -d and -c (delete orphaned objects and create missing OST<br>
&gt; &gt; objects referenced by MDS)<br>
&gt; &gt;<br>
&gt; &gt; After above operations, on clients we were seeing file in red and<br>
&gt; &gt; blinking. Doing a stat came out with an error stating &#39;no such file or<br>
&gt; &gt; directory&#39;.<br>
&gt; &gt;<br>
&gt; &gt; My question is whether the order in which lfsck was run (should lfsck be<br>
&gt; &gt; run multiple times) and the errors we are getting are related or not.<br>
&gt; &gt;<br>
&gt; &gt;<br>
&gt; &gt;<br>
&gt; &gt;<br>
&gt; &gt; --<br>
&gt; &gt; Regards--<br>
&gt; &gt; Rishi Pathak<br>
&gt; &gt; National PARAM Supercomputing Facility<br>
&gt; &gt; Center for Development of Advanced Computing(C-DAC)<br>
&gt; &gt; Pune University Campus,Ganesh Khind Road<br>
&gt; &gt; Pune-Maharastra<br>
&gt; &gt;<br>
&gt; &gt; _______________________________________________<br>
&gt; &gt; Lustre-discuss mailing list<br>
&gt; &gt; <a href="mailto:Lustre-discuss@lists.lustre.org" target="_blank">Lustre-discuss@lists.lustre.org</a><br>
&gt; &gt; <a href="http://lists.lustre.org/mailman/listinfo/lustre-discuss" target="_blank">http://lists.lustre.org/mailman/listinfo/lustre-discuss</a><br>
&gt;<br>
<br>
<br>
--<br>
</div></div><font color="#888888">Bernd Schubert<br>
DataDirect Networks<br>
</font></blockquote></div><br><br clear="all"><br>-- <br>--<br>Wojciech Turek<br><br>Assistant System Manager<br><br>High Performance Computing Service<br>University of Cambridge<br>Email: <a href="mailto:wjt27@cam.ac.uk" target="_blank">wjt27@cam.ac.uk</a><br>

Tel: (+)44 1223 763517 <br>