From dandekar.abhay at gmail.com Sat Aug 9 05:17:23 2014 From: dandekar.abhay at gmail.com (Abhay Dandekar) Date: Sat, 09 Aug 2014 05:17:23 -0000 Subject: [Lustre-devel] Fwd: Lustre configuration failure : lwp-MDT0000: Communicating with 0@lo, operation mds_connect failed with -11. In-Reply-To: References: Message-ID: FWDing ahead to lustre-devel. Requesting some pointers to go ahead. Warm Regards, Abhay Dandekar ---------- Forwarded message ---------- From: Abhay Dandekar Date: Wed, Aug 6, 2014 at 12:18 AM Subject: Lustre configuration failure : lwp-MDT0000: Communicating with 0 at lo, operation mds_connect failed with -11. To: lustre-discuss at lists.lustre.org Hi All, I have come across an lustre installation failure where the MGS is always trying to reach "lo" config instead of configured ethernet. These same steps worked on a different machine, somehow they are failing here. Here are the logs Lustre installation is success with all the packages installed without any error. 0. Lustre version Aug 5 23:07:37 lfs-server kernel: LNet: HW CPU cores: 1, npartitions: 1 Aug 5 23:07:37 lfs-server modprobe: FATAL: Error inserting crc32c_intel (/lib/modules/2.6.32-431.17.1.el6_lustre.x86_64/kernel/arch/x86/crypto/crc32c-intel.ko): No such device Aug 5 23:07:37 lfs-server kernel: alg: No test for crc32 (crc32-table) Aug 5 23:07:37 lfs-server kernel: alg: No test for adler32 (adler32-zlib) Aug 5 23:07:41 lfs-server modprobe: FATAL: Error inserting padlock_sha (/lib/modules/2.6.32-431.17.1.el6_lustre.x86_64/kernel/drivers/crypto/padlock-sha.ko): No such device Aug 5 23:07:41 lfs-server kernel: padlock: VIA PadLock Hash Engine not detected. Aug 5 23:07:45 lfs-server kernel: Lustre: Lustre: Build Version: 2.5.2-RC2--PRISTINE-2.6.32-431.17.1.el6_lustre.x86_64 Aug 5 23:07:45 lfs-server kernel: LNet: Added LNI 192.168.122.50 at tcp [8/256/0/180] Aug 5 23:07:45 lfs-server kernel: LNet: Accept secure, port 988 1. Mkfs [root at lfs-server ~]# mkfs.lustre --fsname=lustre --mgs --mdt --index=0 /dev/sdb Permanent disk data: Target: lustre:MDT0000 Index: 0 Lustre FS: lustre Mount type: ldiskfs Flags: 0x65 (MDT MGS first_time update ) Persistent mount opts: user_xattr,errors=remount-ro Parameters: checking for existing Lustre data: not found device size = 10240MB formatting backing filesystem ldiskfs on /dev/sdb target name lustre:MDT0000 4k blocks 2621440 options -J size=400 -I 512 -i 2048 -q -O dirdata,uninit_bg,^extents,dir_nlink,quota,huge_file,flex_bg -E lazy_journal_init -F mkfs_cmd = mke2fs -j -b 4096 -L lustre:MDT0000 -J size=400 -I 512 -i 2048 -q -O dirdata,uninit_bg,^extents,dir_nlink,quota,huge_file,flex_bg -E lazy_journal_init -F /dev/sdb 2621440 Aug 5 17:16:47 lfs-server kernel: LDISKFS-fs (sdb): mounted filesystem with ordered data mode. quota=on. Opts: Writing CONFIGS/mountdata [root at lfs-server ~]# 2. Mount [root at lfs-server ~]# mount -t lustre /dev/sdb /mnt/mgs Aug 5 17:18:01 lfs-server kernel: LDISKFS-fs (sdb): mounted filesystem with ordered data mode. quota=on. Opts: Aug 5 17:18:01 lfs-server kernel: LDISKFS-fs (sdb): mounted filesystem with ordered data mode. quota=on. Opts: Aug 5 17:18:02 lfs-server kernel: Lustre: ctl-lustre-MDT0000: No data found on store. Initialize space Aug 5 17:18:02 lfs-server kernel: Lustre: lustre-MDT0000: new disk, initializing Aug 5 17:18:02 lfs-server kernel: Lustre: MGS: non-config logname received: params Aug 5 17:18:02 lfs-server kernel: LustreError: 11-0: lustre-MDT0000-lwp-MDT0000: Communicating with 0 at lo, operation mds_connect failed with -11. [root at lfs-server ~]# 3. Unmount [root at lfs-server ~]# umount /dev/sdb Aug 5 17:19:46 lfs-server kernel: Lustre: Failing over lustre-MDT0000 Aug 5 17:19:52 lfs-server kernel: Lustre: 1338:0:(client.c:1908:ptlrpc_expire_one_request()) @@@ Request sent has timed out for slow reply: [sent 1407239386/real 1407239386] req at ffff88003d795c00 x1475596948340888/t0(0) o251->MGC192.168.122.50 at tcp @0 at lo:26/25 lens 224/224 e 0 to 1 dl 1407239392 ref 2 fl Rpc:XN/0/ffffffff rc 0/-1 [root at lfs-server ~]# Aug 5 17:19:53 lfs-server kernel: Lustre: server umount lustre-MDT0000 complete [root at lfs-server ~]# 4. [root at mgs ~]# cat /etc/modprobe.d/lustre.conf options lnet networks=tcp(eth0) [root at mgs ~]# 5.Even the lnet configuration is in place, it does not pick up the required eth0. [root at mgs ~]# lctl dl 0 UP osd-ldiskfs lustre-MDT0000-osd lustre-MDT0000-osd_UUID 8 1 UP mgs MGS MGS 5 2 UP mgc MGC192.168.122.50 at tcp c6ea84c0-b3b2-9d25-8126-32d85956ae4d 5 3 UP mds MDS MDS_uuid 3 4 UP lod lustre-MDT0000-mdtlov lustre-MDT0000-mdtlov_UUID 4 5 UP mdt lustre-MDT0000 lustre-MDT0000_UUID 5 6 UP mdd lustre-MDD0000 lustre-MDD0000_UUID 4 7 UP qmt lustre-QMT0000 lustre-QMT0000_UUID 4 8 UP lwp lustre-MDT0000-lwp-MDT0000 lustre-MDT0000-lwp-MDT0000_UUID 5 [root at mgs ~]# Any pointers to go ahead ?? Warm Regards, Abhay Dandekar -------------- next part -------------- An HTML attachment was scrubbed... URL: From aayush.agrawal at calsoftinc.com Mon Aug 18 10:27:25 2014 From: aayush.agrawal at calsoftinc.com (aayush agrawal) Date: Mon, 18 Aug 2014 15:57:25 +0530 Subject: [Lustre-devel] Full stripe write in RAID6 Message-ID: <53F1D50D.7080805@calsoftinc.com> Hi, I am using lustre version: 2.5.0 and corresponding kernel 2.6.32-358. Apart from default patches which comes with lustre 2.5.0, I applied below patches in kernel 2.6.32-358. raid5-configurable-cachesize-rhel6.patch raid5-large-io-rhel5.patch raid5-stats-rhel6.patch raid5-zerocopy-rhel6.patch raid5-mmp-unplug-dev-rhel6.patch raid5-mmp-unplug-dev.patch raid5-maxsectors-rhel5.patch raid5-stripe-by-stripe-handling-rhel6.patch I have taken all above patches from below link: https://github.com/Xyratex/lustre-stable/tree/b_neo_1.4.0/lustre/kernel_patches/patches My question is: If I am writing entire stripe then whether RAID6 md driver need to read any of the blocks from underlying device? I am asking this question on lustre mailing list because I have seen that lustre community has changed RAID driver a lot. I have created RAID6 device with default (512K) chunk size with total 6 RAID devices. cat /sys/block/md127/queue/optimal_io_size =>2097152 I believe this is full stripe (512K * 4 data disks). If I write 2MB data, I am expected to dirty entire stripe hence what I believe I need not require to read either any of the data block or parity blocks. Thus avoiding RAID6 penalties. Whether md/raid driver supports full stripe writes by avoiding RAID 6 penalties? I also expected 6 disks will receive 512K writes each. (4 data disk+ 2 parity disks). If I do IO directly on block device /dev/md127, I do observe reads happening on md device and underlying raid devices as well. #mdstat o/p: md127 : active raid6 sdah1[5] sdai1[4] sdaj1[3] sdcg1[2] sdch1[1] sdci1[0] 41926656 blocks super 1.2 level 6, 512k chunk, algorithm 2 [6/6] [UUUUUU] # raw -qa /dev/raw/raw1: bound to major 9, minor 127 #time (dd if=/dev/zero of=/dev/raw/raw1 bs=2M count=1 && sync) (also tried with of=/dev/md127 oflag=direct but the same results.) # iostat shows: Device: tps Blk_read/s Blk_wrtn/s Blk_read Blk_wrtn sdaj1 7.00 0.00 205.20 0 1026 sdai1 6.20 0.00 205.20 0 1026 sdah1 9.80 0.00 246.80 0 1234 sdcg1 6.80 0.00 205.20 0 1026 sdci1 9.60 0.00 246.80 0 1234 sdch1 6.80 0.00 205.20 0 1026 md127 0.80 0.00 819.20 0 4096 I assume if I perform writes in multiples of "optimal_io_size" I would be doing full stripe writes thus avoiding reads. But unfortunately with two 2M writes, I do see reads happening for some these drives. Same case for count=4 or 6 (equal to data disks or total disks). # time (dd if=/dev/zero of=/dev/raw/raw1 bs=2M count=2 && sync) Device: tps Blk_read/s Blk_wrtn/s Blk_read Blk_wrtn sdaj1 13.40 204.80 410.00 1024 2050 sdai1 11.20 0.00 410.00 0 2050 sdah1 15.80 0.00 464.40 0 2322 sdcg1 13.20 204.80 410.00 1024 2050 sdci1 16.60 0.00 464.40 0 2322 sdch1 12.40 192.00 410.00 960 2050 md127 1.60 0.00 1638.40 0 8192 I believe RAID6 penalties will exist if it's a random write, but in case of seq. write, whether they will still exist in some other form in Linux md/raid driver? My aim is to maximize RAID6 Write IO rate with sequential Writes withoutRAID6 penalties. Rectify me wherever my assumptions are wrong. Let me know if any other configuration param (for block device or md device) is required to achieve the same. Thanks, Aayush -------------- next part -------------- An HTML attachment was scrubbed... URL: From pratik.rupala at calsoftinc.com Mon Aug 18 11:49:07 2014 From: pratik.rupala at calsoftinc.com (Pratik Rupala) Date: Mon, 18 Aug 2014 17:19:07 +0530 Subject: [Lustre-devel] Question regarding lustre patches Message-ID: <53F1E833.1040100@calsoftinc.com> Hi, In the older version of lustre, I can see a patch named "*raid5-merge-ios-rhel5.patch*" which is related to RAID performance improvement for RHEL5. As per my understanding, it accumulates the bios in a single make_request function of RAID layer and sends them collectively by generic_make_request instead of sending them separately per stripe basis. By this means, performance of RAID could be improved. But this patch is not present in newer version of lustre and especially for RHEL6. So, is it possible to port that patch for RHEL 6, provided RAID architecture is changed significantly and handling of stripe has been made asynchronous in RHEL 6 as compared to synchronous handling of stripe in RHEL 5? And even if it can be ported for RHEL6, will it give advantage as it was giving in RHEL 5 with different RAID architecture? Regards, Pratik -------------- next part -------------- An HTML attachment was scrubbed... URL: From xanthornus at gmail.com Fri Aug 29 13:24:57 2014 From: xanthornus at gmail.com (Rituparna Chakraborty) Date: Fri, 29 Aug 2014 18:54:57 +0530 Subject: [Lustre-devel] request to work on the project Message-ID: I would like to work on ioctl() cleanups project -------------- next part -------------- An HTML attachment was scrubbed... URL: