From vuhuong at mellanox.com Wed Mar 2 19:57:06 2011 From: vuhuong at mellanox.com (Vu Pham) Date: Wed, 02 Mar 2011 11:57:06 -0800 Subject: [Lustre-devel] system crashes mounting mds Message-ID: <4D6EA112.8040400@mellanox.com> Hi, I got system crash with message "BUG: scheduling while atomic: ll_mgs_01/0xffff8103/11347" after mounting lustre Here is the steps that I did: $ mkfs.lustre --fsname=lustre --reformat --mgs --mdt /dev/sdc $ mount -t lustre /dev/sdc /tmp/lustre_mgs Here is the stack dump: ldiskfs created from ext3-2.6-rhel5 kjournald starting. Commit interval 5 seconds LDISKFS FS on sdc, internal journal LDISKFS-fs: mounted filesystem with ordered data mode. Lustre: OBD class driver, http://www.lustre.org/ Lustre: Lustre Version: 1.8.5 Lustre: Build Version: 1.8.5-20101117053234-PRISTINE-2.6.18-194.17.1.el5_lustre.1.8.5 Lustre: Added LNI 10.4.57.8 at tcp [8/256/0/180] Lustre: Accept secure, port 988 Lustre: Lustre Client File System; http://www.lustre.org/ kjournald starting. Commit interval 5 seconds LDISKFS FS on sdc, internal journal LDISKFS-fs: mounted filesystem with ordered data mode. kjournald starting. Commit interval 5 seconds LDISKFS FS on sdc, internal journal LDISKFS-fs: mounted filesystem with ordered data mode. Lustre: MGS MGS started Lustre: MGC10.4.57.8 at tcp: Reactivating import Lustre: MGS: Logs for fs lustre were removed by user request. All servers must be restarted in order to regenerate the logs. BUG: scheduling while atomic: ll_mgs_01/0xffff8103/11347 Call Trace: [] __sched_text_start+0x7d/0xbd6 [] :scsi_mod:scsi_done+0x0/0x18 [] __mod_timer+0x100/0x10f [] do_gettimeofday+0x40/0x90 [] getnstimeofday+0x10/0x28 [] sync_buffer+0x0/0x3f [] io_schedule+0x3f/0x67 [] sync_buffer+0x3b/0x3f [] __wait_on_bit+0x40/0x6e [] sync_buffer+0x0/0x3f [] out_of_line_wait_on_bit+0x6c/0x78 [] wake_bit_function+0x0/0x23 [] :ldiskfs:bh_submit_read+0x58/0x70 [] :ldiskfs:read_block_bitmap+0xc8/0x1c0 [] :ldiskfs:ldiskfs_new_blocks_old+0x1df/0x750 [] :ldiskfs:ldiskfs_get_blocks_handle+0x596/0xd30 [] :ldiskfs:ldiskfs_get_blocks_handle+0x11a/0xd30 [] :ldiskfs:ldiskfs_get_blocks_handle+0x11a/0xd30 [] __find_get_block+0x15c/0x16c [] :ldiskfs:ldiskfs_getblk+0xea/0x320 [] :jbd:start_this_handle+0x341/0x3ed [] __getblk+0x25/0x236 [] :ldiskfs:ldiskfs_bread+0x11/0x80 [] :jbd:journal_start+0xd3/0x107 [] :fsfilt_ldiskfs:fsfilt_ldiskfs_write_record+0x1cd/0x4b0 [] do_lookup+0x65/0x1e6 [] :obdclass:llog_lvfs_write_blob+0x119/0x440 [] :obdclass:llog_lvfs_write_rec+0xb1f/0xda0 [] file_move+0x36/0x44 [] dput+0x2c/0x113 [] :mgs:record_lcfg+0x38e/0x4c0 [] __d_lookup+0xb0/0xff [] :mgs:record_marker+0x83a/0xa30 [] mntput_no_expire+0x19/0x89 [] :mgs:mgs_write_log_lov+0x37b/0xf80 [] snprintf+0x44/0x4c [] :lvfs:pop_ctxt+0x290/0x370 [] :obdclass:__llog_ctxt_put+0x26/0x150 [] :mgs:__mgs_write_log_mdt+0x2b3/0x5d0 [] :mgs:mgs_write_log_target+0xb5f/0x21e0 [] :ptlrpc:ldlm_completion_ast+0x0/0x880 [] :mgs:mgs_handle+0xf09/0x16c0 [] :ptlrpc:ptlrpc_server_handle_request+0x97a/0xdf0 [] :ptlrpc:ptlrpc_wait_event+0x2d8/0x310 [] __wake_up_common+0x3e/0x68 [] :ptlrpc:ptlrpc_main+0xf37/0x10f0 [] child_rip+0xa/0x11 [] :ptlrpc:ptlrpc_main+0x0/0x10f0 [] child_rip+0x0/0x11 Unable to handle kernel paging request at ffffffffe2cf5e40 RIP: [] __sched_text_start+0x72e/0xbd6 PGD 203067 PUD 10af48067 PMD 0 Oops: 0000 [1] SMP last sysfs file: /class/misc/obd/dev CPU 2 Modules linked in: mds(U) fsfilt_ldiskfs(U) mgs(U) mgc(U) lustre(U) lov(U) mdc(U) lquota(U) osc(U) ksocklnd(U) ptlrpc(U) obdclass(U) lnet(U) lvfs(U) libcfs(U) ldiskfs(U) crc16(U) mlx4_fcoib(U) mlx4_fc(U) libfc(U) scsi_transport_fc(U) netconsole(U) nfs(U) fscache(U) nfsd(U) exportfs(U) nfs_acl(U) auth_rpcgss(U) autofs4(U) rdma_ucm(U) rdma_cm(U) ib_cm(U) iw_cm(U) ib_sa(U) ib_addr(U) ib_uverbs(U) ib_umad(U) mlx4_ib(U) ib_mad(U) ib_core(U) mlx4_en(U) mlx4_core(U) hidp(U) l2cap(U) bluetooth(U) lockd(U) sunrpc(U) ipv6(U) xfrm_nalgo(U) crypto_api(U) vfat(U) fat(U) loop(U) dm_mirror(U) dm_multipath(U) scsi_dh(U) video(U) backlight(U) sbs(U) power_meter(U) hwmon(U) i2c_ec(U) i2c_core(U) dell_wmi(U) wmi(U) button(U) battery(U) asus_acpi(U) acpi_memhotplug(U) ac(U) parport_pc(U) lp(U) parport(U) sr_mod(U) cdrom(U) sg(U) hpilo(U) bnx2(U) serio_raw(U) pcspkr(U) dm_raid45(U) dm_message(U) dm_region_hash(U) dm_log(U) dm_mod(U) dm_mem_cache(U) ata_piix(U) libata(U) shpchp(U) cciss(U) sd_mod (U) scsi_mod(U) ext3(U) jbd(U) uhci_hcd(U) ohci_hcd(U) ehci_hcd(U) Pid: 11347, comm: ll_mgs_01 Tainted: G 2.6.18-194.17.1.el5_lustre.1.8.5 #1 RIP: 0010:[] [] __sched_text_start+0x72e/0xbd6 RSP: 0000:ffff8102fd9130b0 EFLAGS: 00010083 RAX: ffffffff80441380 RBX: ffff8102f98aa7a0 RCX: 0000031f0fbe6ec8 RDX: 000000000c520680 RSI: ffff8102f98aa7a0 RDI: ffff8102f98aa7a0 RBP: ffff8102fd913170 R08: 00000000000000a0 R09: 0000000000000000 R10: ffffffff80015504 R11: 0000000000000000 R12: ffff8102f98aa7d8 R13: ffff81010af445f8 R14: 0000000000000002 R15: ffff81000100caa0 FS: 00002af059d15230(0000) GS:ffff81010af994c0(0000) knlGS:0000000000000000 CS: 0010 DS: 0000 ES: 0000 CR0: 000000008005003b CR2: ffffffffe2cf5e40 CR3: 000000061d85d000 CR4: 00000000000006e0 Process ll_mgs_01 (pid: 11347, threadinfo ffff8102fd912000, task ffff8102f98aa7a0) Stack: ffff81031d634c88 ffffffff880765a6 ffff81031d634c80 0000000000000004 ffff8102f98aa7a0 ffff8102f98aa7a0 0000031f0fc500bd 000000000f14b06c ffff8102f98aa990 000000021d634c80 ffff810611200000 ffffffff8006e1d7 Call Trace: [] :scsi_mod:scsi_done+0x0/0x18 [] do_gettimeofday+0x40/0x90 [] getnstimeofday+0x10/0x28 [] sync_buffer+0x0/0x3f [] io_schedule+0x3f/0x67 [] sync_buffer+0x3b/0x3f [] __wait_on_bit+0x40/0x6e [] sync_buffer+0x0/0x3f [] out_of_line_wait_on_bit+0x6c/0x78 [] wake_bit_function+0x0/0x23 [] :ldiskfs:bh_submit_read+0x58/0x70 [] :ldiskfs:read_block_bitmap+0xc8/0x1c0 [] :ldiskfs:ldiskfs_new_blocks_old+0x1df/0x750 [] :ldiskfs:ldiskfs_get_blocks_handle+0x596/0xd30 [] :ldiskfs:ldiskfs_get_blocks_handle+0x11a/0xd30 [] :ldiskfs:ldiskfs_get_blocks_handle+0x11a/0xd30 [] __find_get_block+0x15c/0x16c [] :ldiskfs:ldiskfs_getblk+0xea/0x320 [] :jbd:start_this_handle+0x341/0x3ed [] __getblk+0x25/0x236 [] :ldiskfs:ldiskfs_bread+0x11/0x80 [] :jbd:journal_start+0xd3/0x107 [] :fsfilt_ldiskfs:fsfilt_ldiskfs_write_record+0x1cd/0x4b0 [] do_lookup+0x65/0x1e6 [] :obdclass:llog_lvfs_write_blob+0x119/0x440 [] :obdclass:llog_lvfs_write_rec+0xb1f/0xda0 [] file_move+0x36/0x44 [] dput+0x2c/0x113 [] :mgs:record_lcfg+0x38e/0x4c0 [] __d_lookup+0xb0/0xff [] :mgs:record_marker+0x83a/0xa30 [] mntput_no_expire+0x19/0x89 [] :mgs:mgs_write_log_lov+0x37b/0xf80 [] snprintf+0x44/0x4c [] :lvfs:pop_ctxt+0x290/0x370 [] :obdclass:__llog_ctxt_put+0x26/0x150 [] :mgs:__mgs_write_log_mdt+0x2b3/0x5d0 [] :mgs:mgs_write_log_target+0xb5f/0x21e0 [] :ptlrpc:ldlm_completion_ast+0x0/0x880 [] :mgs:mgs_handle+0xf09/0x16c0 [] :ptlrpc:ptlrpc_server_handle_request+0x97a/0xdf0 [] :ptlrpc:ptlrpc_wait_event+0x2d8/0x310 [] __wake_up_common+0x3e/0x68 [] :ptlrpc:ptlrpc_main+0xf37/0x10f0 [] child_rip+0xa/0x11 [] :ptlrpc:ptlrpc_main+0x0/0x10f0 [] child_rip+0x0/0x11 Code: 48 8b 14 d5 40 2a 3f 80 48 03 42 08 31 d2 c7 40 08 01 00 00 RIP [] __sched_text_start+0x72e/0xbd6 RSP CR2: ffffffffe2cf5e40 <0>Kernel panic - not syncing: Fatal exception By the way, I also try the same setup steps on different device ie. /dev/cciss/c0d0p6 and it is fine. I'm writing scsi lld driver FCoIB, sdc is scsi device (ie. FC lun) seen/controlled by FCoIB driver, I can mount filesystems ext2/ext3/reiserfs... and run normal I/O on sdc without problem. Could anyone help/shed some lights on what the problem is? thanks, -vu From adilger at whamcloud.com Wed Mar 2 21:23:52 2011 From: adilger at whamcloud.com (Andreas Dilger) Date: Wed, 2 Mar 2011 14:23:52 -0700 Subject: [Lustre-devel] system crashes mounting mds In-Reply-To: <4D6EA112.8040400@mellanox.com> References: <4D6EA112.8040400@mellanox.com> Message-ID: <2938460B-B801-4692-8C79-2AC4A7B29C27@whamcloud.com> On 2011-03-02, at 12:57 PM, Vu Pham wrote: > I got system crash with message "BUG: scheduling while atomic: > Lustre: Lustre Version: 1.8.5 > Lustre: Build Version: 1.8.5-20101117053234-PRISTINE-2.6.18-194.17.1.el5_lustre.1.8.5 > ll_mgs_01/0xffff8103/11347" after mounting lustre This is often caused by a stack overflow. Looking at the stack trace, it _shouldn't_ be atomic in that context due to Lustre (submitting a block IO) so I suspect that the "preempt_count" in the tast struct is corrupted or similar. > Here is the stack dump: > > BUG: scheduling while atomic: ll_mgs_01/0xffff8103/11347 > > Call Trace: > [] __sched_text_start+0x7d/0xbd6 > [] :scsi_mod:scsi_done+0x0/0x18 > [] __mod_timer+0x100/0x10f > [] do_gettimeofday+0x40/0x90 > [] getnstimeofday+0x10/0x28 > [] sync_buffer+0x0/0x3f > [] io_schedule+0x3f/0x67 > [] sync_buffer+0x3b/0x3f > [] __wait_on_bit+0x40/0x6e > [] sync_buffer+0x0/0x3f > [] out_of_line_wait_on_bit+0x6c/0x78 > [] wake_bit_function+0x0/0x23 > [] :ldiskfs:bh_submit_read+0x58/0x70 > [] :ldiskfs:read_block_bitmap+0xc8/0x1c0 > [] :ldiskfs:ldiskfs_new_blocks_old+0x1df/0x750 > [] :ldiskfs:ldiskfs_get_blocks_handle+0x596/0xd30 > [] :ldiskfs:ldiskfs_get_blocks_handle+0x11a/0xd30 > [] :ldiskfs:ldiskfs_get_blocks_handle+0x11a/0xd30 > [] __find_get_block+0x15c/0x16c > [] :ldiskfs:ldiskfs_getblk+0xea/0x320 > [] :jbd:start_this_handle+0x341/0x3ed > [] __getblk+0x25/0x236 > [] :ldiskfs:ldiskfs_bread+0x11/0x80 > [] :jbd:journal_start+0xd3/0x107 > [] :fsfilt_ldiskfs:fsfilt_ldiskfs_write_record+0x1cd/0x4b0 > [] do_lookup+0x65/0x1e6 > [] :obdclass:llog_lvfs_write_blob+0x119/0x440 > [] :obdclass:llog_lvfs_write_rec+0xb1f/0xda0 > [] file_move+0x36/0x44 > [] dput+0x2c/0x113 > [] :mgs:record_lcfg+0x38e/0x4c0 > [] __d_lookup+0xb0/0xff > [] :mgs:record_marker+0x83a/0xa30 > [] mntput_no_expire+0x19/0x89 > [] :mgs:mgs_write_log_lov+0x37b/0xf80 > [] snprintf+0x44/0x4c > [] :lvfs:pop_ctxt+0x290/0x370 > [] :obdclass:__llog_ctxt_put+0x26/0x150 > [] :mgs:__mgs_write_log_mdt+0x2b3/0x5d0 > [] :mgs:mgs_write_log_target+0xb5f/0x21e0 > [] :ptlrpc:ldlm_completion_ast+0x0/0x880 > [] :mgs:mgs_handle+0xf09/0x16c0 > [] :ptlrpc:ptlrpc_server_handle_request+0x97a/0xdf0 > [] :ptlrpc:ptlrpc_wait_event+0x2d8/0x310 > [] __wake_up_common+0x3e/0x68 > [] :ptlrpc:ptlrpc_main+0xf37/0x10f0 > [] child_rip+0xa/0x11 > [] :ptlrpc:ptlrpc_main+0x0/0x10f0 > [] child_rip+0x0/0x11 > > > By the way, I also try the same setup steps on different device ie. /dev/cciss/c0d0p6 and it is fine. > > I'm writing scsi lld driver FCoIB, sdc is scsi device (ie. FC lun) seen/controlled by FCoIB driver, I can mount filesystems ext2/ext3/reiserfs... and run normal I/O on sdc without problem. Those filesystems use far less stack - Lustre is using a bunch of extra stack on top of ext4 (i.e. everything on top of "fsfilt_ldiskfs_write_record()" is on top of the stack usage of the local filesystem. > Could anyone help/shed some lights on what the problem is? Cheers, Andreas -- Andreas Dilger Principal Engineer Whamcloud, Inc. From vuhuong at mellanox.com Wed Mar 2 23:07:21 2011 From: vuhuong at mellanox.com (Vu Pham) Date: Wed, 02 Mar 2011 15:07:21 -0800 Subject: [Lustre-devel] system crashes mounting mds In-Reply-To: <2938460B-B801-4692-8C79-2AC4A7B29C27@whamcloud.com> References: <4D6EA112.8040400@mellanox.com> <2938460B-B801-4692-8C79-2AC4A7B29C27@whamcloud.com> Message-ID: <4D6ECDA9.3090500@mellanox.com> Andreas Dilger wrote: > On 2011-03-02, at 12:57 PM, Vu Pham wrote: >> I got system crash with message "BUG: scheduling while atomic: >> Lustre: Lustre Version: 1.8.5 >> Lustre: Build Version: > 1.8.5-20101117053234-PRISTINE-2.6.18-194.17.1.el5_lustre.1.8.5 >> ll_mgs_01/0xffff8103/11347" after mounting lustre > > This is often caused by a stack overflow. > > Looking at the stack trace, it _shouldn't_ be atomic in that context due > to Lustre (submitting a block IO) so I suspect that the "preempt_count" > in the tast struct is corrupted or similar. > >> Here is the stack dump: >> >> BUG: scheduling while atomic: ll_mgs_01/0xffff8103/11347 >> >> Call Trace: >> [] __sched_text_start+0x7d/0xbd6 >> [] :scsi_mod:scsi_done+0x0/0x18 >> [] __mod_timer+0x100/0x10f >> [] do_gettimeofday+0x40/0x90 >> [] getnstimeofday+0x10/0x28 >> [] sync_buffer+0x0/0x3f >> [] io_schedule+0x3f/0x67 >> [] sync_buffer+0x3b/0x3f >> [] __wait_on_bit+0x40/0x6e >> [] sync_buffer+0x0/0x3f >> [] out_of_line_wait_on_bit+0x6c/0x78 >> [] wake_bit_function+0x0/0x23 >> [] :ldiskfs:bh_submit_read+0x58/0x70 >> [] :ldiskfs:read_block_bitmap+0xc8/0x1c0 >> [] :ldiskfs:ldiskfs_new_blocks_old+0x1df/0x750 >> [] :ldiskfs:ldiskfs_get_blocks_handle+0x596/0xd30 >> [] :ldiskfs:ldiskfs_get_blocks_handle+0x11a/0xd30 >> [] :ldiskfs:ldiskfs_get_blocks_handle+0x11a/0xd30 >> [] __find_get_block+0x15c/0x16c >> [] :ldiskfs:ldiskfs_getblk+0xea/0x320 >> [] :jbd:start_this_handle+0x341/0x3ed >> [] __getblk+0x25/0x236 >> [] :ldiskfs:ldiskfs_bread+0x11/0x80 >> [] :jbd:journal_start+0xd3/0x107 >> [] > :fsfilt_ldiskfs:fsfilt_ldiskfs_write_record+0x1cd/0x4b0 >> [] do_lookup+0x65/0x1e6 >> [] :obdclass:llog_lvfs_write_blob+0x119/0x440 >> [] :obdclass:llog_lvfs_write_rec+0xb1f/0xda0 >> [] file_move+0x36/0x44 >> [] dput+0x2c/0x113 >> [] :mgs:record_lcfg+0x38e/0x4c0 >> [] __d_lookup+0xb0/0xff >> [] :mgs:record_marker+0x83a/0xa30 >> [] mntput_no_expire+0x19/0x89 >> [] :mgs:mgs_write_log_lov+0x37b/0xf80 >> [] snprintf+0x44/0x4c >> [] :lvfs:pop_ctxt+0x290/0x370 >> [] :obdclass:__llog_ctxt_put+0x26/0x150 >> [] :mgs:__mgs_write_log_mdt+0x2b3/0x5d0 >> [] :mgs:mgs_write_log_target+0xb5f/0x21e0 >> [] :ptlrpc:ldlm_completion_ast+0x0/0x880 >> [] :mgs:mgs_handle+0xf09/0x16c0 >> [] :ptlrpc:ptlrpc_server_handle_request+0x97a/0xdf0 >> [] :ptlrpc:ptlrpc_wait_event+0x2d8/0x310 >> [] __wake_up_common+0x3e/0x68 >> [] :ptlrpc:ptlrpc_main+0xf37/0x10f0 >> [] child_rip+0xa/0x11 >> [] :ptlrpc:ptlrpc_main+0x0/0x10f0 >> [] child_rip+0x0/0x11 >> I get the above stack trace when I called scsi_done() in workqueue context. Originally calling scsi_done() in irq context, I get below stack trace. Unable to handle kernel paging request at 00000000031b1e40 RIP: [] task_rq_lock+0x29/0x6f PGD 61d9b3067 PUD 61d7a3067 PMD 0 Oops: 0000 [1] SMP last sysfs file: /class/misc/obd/dev CPU 2 Modules linked in: mds(U) fsfilt_ldiskfs(U) mgs(U) mgc(U) ldiskfs(U) crc16(U) lustre(U) lov(U) mdc(U) lquota(U) osc(U) ksocklnd(U) ptlrpc(U) obdclass(U) lnet(U) lvfs(U) libcfs(U) mlx4_fcoib(U) mlx4_fc(U) libfc(U) scsi_transport_fc(U) netconsole(U) nfs(U) fscache(U) nfsd(U) exportfs(U) nfs_acl(U) auth_rpcgss(U) autofs4(U) rdma_ucm(U) rdma_cm(U) ib_cm(U) iw_cm(U) ib_sa(U) ib_addr(U) ib_uverbs(U) ib_umad(U) mlx4_ib(U) ib_mad(U) ib_core(U) mlx4_en(U) mlx4_core(U) hidp(U) l2cap(U) bluetooth(U) lockd(U) sunrpc(U) ipv6(U) xfrm_nalgo(U) crypto_api(U) vfat(U) fat(U) loop(U) dm_mirror(U) dm_multipath(U) scsi_dh(U) video(U) backlight(U) sbs(U) power_meter(U) hwmon(U) i2c_ec(U) i2c_core(U) dell_wmi(U) wmi(U) button(U) battery(U) asus_acpi(U) acpi_memhotplug(U) ac(U) parport_pc(U) lp(U) parport(U) sr_mod(U) cdrom(U) sg(U) hpilo(U) serio_raw(U) pcspkr(U) bnx2(U) dm_raid45(U) dm_message(U) dm_region_hash(U) dm_log(U) dm_mod(U) dm_mem_cache(U) ata_piix(U) libata(U) shpchp(U) cciss(U) sd_mod (U) scsi_mod(U) ext3(U) jbd(U) uhci_hcd(U) ohci_hcd(U) ehci_hcd(U) Pid: 0, comm: swapper Tainted: G 2.6.18-194.17.1.el5_lustre.1.8.5 #1 RIP: 0010:[] [] task_rq_lock+0x29/0x6f RSP: 0000:ffff81010aff3c30 EFLAGS: 00010086 RAX: 00000000105b7e80 RBX: ffffffff8043f420 RCX: ffff81010aff3d90 RDX: 0000000000000000 RSI: ffff81010aff3cb8 RDI: ffff8102f10ac040 RBP: ffff81010aff3c50 R08: ffff8102ebe214f0 R09: ffff81010aff3f10 R10: ffffffff8003da79 R11: ffffffff80041fc2 R12: ffffffff8043f420 R13: ffff81010aff3cb8 R14: ffff8102f10ac040 R15: ffff81010aff3d90 FS: 0000000000000000(0000) GS:ffff81010af994c0(0000) knlGS:0000000000000000 CS: 0010 DS: 0018 ES: 0018 CR0: 000000008005003b CR2: 00000000031b1e40 CR3: 0000000619237000 CR4: 00000000000006e0 Process swapper (pid: 0, threadinfo ffff81010afee000, task ffff81032a957860) Stack: 0000000000000003 0000000000000001 ffff8102f10ac040 ffff81010af38610 ffff81010aff3cf0 ffffffff80046a5c 00000000803f85a0 0000000000000030 ffff81010afeff00 0000000000020000 0000000000020000 00000000000456c4 Call Trace: [] try_to_wake_up+0x27/0x484 [] start_secondary+0x498/0x4a7 [] start_secondary+0x498/0x4a7 [] autoremove_wake_function+0x9/0x2e [] __wake_up_common+0x3e/0x68 [] __wake_up+0x38/0x4f [] __wake_up_bit+0x28/0x2d [] end_buffer_read_sync+0x1c/0x22 [] end_bio_bh_io_sync+0x2f/0x3b [] __end_that_request_first+0x23c/0x5bf [] show_trace+0x34/0x47 [] :scsi_mod:scsi_end_request+0x27/0xcd [] :scsi_mod:scsi_io_completion+0x14e/0x324 [] :mlx4_fc:mfc_cq_clean+0x4f/0x84 [] :sd_mod:sd_rw_intr+0x25a/0x294 [] :scsi_mod:scsi_device_unbusy+0x67/0x81 [] blk_done_softirq+0x5f/0x6d [] __do_softirq+0x89/0x133 [] call_softirq+0x1c/0x28 [] do_softirq+0x2c/0x85 [] do_IRQ+0xec/0xf5 [] ret_from_intr+0x0/0xa [] :ptlrpc:request_out_callback+0x0/0x1b0 [] acpi_processor_idle_simple+0x17d/0x30e [] acpi_processor_idle_simple+0x6c/0x30e [] acpi_processor_idle_simple+0x0/0x30e [] acpi_processor_idle_simple+0x0/0x30e [] cpu_idle+0x95/0xb8 [] start_secondary+0x498/0x4a7 Code: 48 8b 04 c5 40 2a 3f 80 4c 03 60 08 4c 89 e7 e8 52 83 fd ff RIP [] task_rq_lock+0x29/0x6f RSP CR2: 00000000031b1e40 <0>Kernel panic - not syncing: Fatal exception Have you seen this problem with other scsi devices on other transport (sata, fc,...)? From jeremy.filizetti at gmail.com Fri Mar 4 04:48:57 2011 From: jeremy.filizetti at gmail.com (Jeremy Filizetti) Date: Thu, 03 Mar 2011 23:48:57 -0500 Subject: [Lustre-devel] lustre 1.8+ issues with automounter Message-ID: <4D706F39.2070807@gmail.com> Ever since we moved from Lustre 1.6.6 to 1.8 I've seen issues with using the automounter and Lustre. I've finally got around to looking at what the issue is, but I'm not quite sure what the correct way to resolve it is. I think the issue will remain in 2.0+ but I didn't look closely at the code. The issue is that lov_connect which calls lov_connect_obd is an asynchronous connect that does not wait for all OSCs to be connected before returning. In the end lustre_fill_super can return before all OSCs have been set active so any file operations that caused the automount may return an error. Many lov functions check to make sure the lov_tgt_desc ltd_active flag is 1 or return -EIO. The following patch handles things correctly by waiting until all OSC's that are set to be activated are active before returning from filling the super block. There are a few problems that I'm not sure of what the expected results are with Lustre. For example if an OST has not been mounted the client will attempt to connect and end up returning -ENODEV and setting the import_state as LUSTRE_IMP_DISCON. Without the patch the client mounts immediately even though the OSC is unavailable, with it the mount would not return until the user kills the process, the OBD is set inactive, or the state changes. To provide the same functionality an extra condition would need to be added to the l_wait_event condition to monitor the import state is not connecting. However if I do that, I'm not sure things handle failover nodes correctly. So what I'm wondering is what are the expected actions for the different conditions of OSTs. Thanks, Jeremy diff --git a/lustre/include/obd.h b/lustre/include/obd.h index e89805d..3046a5c 100644 --- a/lustre/include/obd.h +++ b/lustre/include/obd.h @@ -754,6 +754,8 @@ struct lov_tgt_desc { unsigned long ltd_active:1,/* is this target up for requests */ ltd_activate:1,/* should this target be activated */ ltd_reap:1; /* should this target be deleted */ + cfs_waitq_t ltd_started; /* waitqueue to notify tgt has been fully started + * so IO can start */ }; /* Pool metadata */ @@ -942,6 +944,8 @@ enum obd_notify_event { OBD_NOTIFY_ACTIVE, /* Device deactivated */ OBD_NOTIFY_INACTIVE, + /* Device disconnected */ + OBD_NOTIFY_DISCON, /* Connect data for import were changed */ OBD_NOTIFY_OCD, /* Sync request */ diff --git a/lustre/lov/lov_obd.c b/lustre/lov/lov_obd.c index 8b2d848..ff4a04a 100644 --- a/lustre/lov/lov_obd.c +++ b/lustre/lov/lov_obd.c @@ -222,7 +222,33 @@ static int lov_notify(struct obd_device *obd, struct obd_device *watched, } /* active event should be pass lov target index as data */ data = &rc; - } + } else if (ev == OBD_NOTIFY_DISCON) { + struct lov_tgt_desc *tgt; + struct lov_obd *lov = &obd->u.lov; + int i; + + LASSERT(watched); + if (strcmp(watched->obd_type->typ_name, LUSTRE_OSC_NAME)) { + CERROR("unexpected notification of %s %s!\n", + watched->obd_type->typ_name, + watched->obd_name); + RETURN(-EINVAL); + } + + obd_getref(obd); + for (i = 0; i < lov->desc.ld_tgt_count; i++) { + tgt = lov->lov_tgts[i]; + if (!tgt || !tgt->ltd_exp) + continue; + + if (obd_uuid_equals(&watched->u.cli.cl_target_uuid, &tgt->ltd_uuid)) { + cfs_waitq_signal(&lov->lov_tgts[i]->ltd_started); + data = &i; + break; + } + } + obd_putref(obd); + } /* Pass the notification up the chain. */ if (watched) { @@ -424,6 +450,27 @@ static int lov_connect(struct lustre_handle *conn, struct obd_device *obd, obd->obd_name, rc); } } + + /* Wait for all the connections to complete before returning so that all + * obds are set active that should be. Otherwise IO that happens immediately + * after mount could (autofs) could glimpse or touch objects before the connecction + * is established */ + for (i = 0; i < lov->desc.ld_tgt_count; i++) { + struct l_wait_info lwi = { 0 }; + + tgt = lov->lov_tgts[i]; + if (!tgt || !tgt->ltd_exp || obd_uuid_empty(&tgt->ltd_uuid)) + continue; + + if (tgt->ltd_activate == tgt->ltd_active) + continue; + + CDEBUG(D_CONFIG, "Target %s activate/active %d/%d, waiting on state change\n", + tgt->ltd_obd->obd_name, tgt->ltd_activate, tgt->ltd_active); + + l_wait_event(tgt->ltd_started, tgt->ltd_activate == tgt->ltd_active || + tgt->ltd_obd->u.cli.cl_import->imp_deactive, &lwi); + } obd_putref(obd); RETURN(0); @@ -445,6 +492,9 @@ static int lov_disconnect_obd(struct obd_device *obd, struct lov_tgt_desc *tgt) tgt->ltd_active = 0; lov->desc.ld_active_tgt_count--; tgt->ltd_exp->exp_obd->obd_inactive = 1; + + /* If state change wake up wait queue */ + cfs_waitq_signal(&tgt->ltd_started); } lov_proc_dir = lprocfs_srch(obd->obd_proc_entry, "target_obds"); @@ -582,6 +632,9 @@ static int lov_set_osc_active(struct obd_device *obd, struct obd_uuid *uuid, lov->lov_tgts[i]->ltd_qos.ltq_penalty = 0; out: + if (i >= 0) + cfs_waitq_signal(&lov->lov_tgts[i]->ltd_started); + obd_putref(obd); RETURN(i); } @@ -673,6 +726,8 @@ static int lov_add_target(struct obd_device *obd, struct obd_uuid *uuidp, if (index >= lov->desc.ld_tgt_count) lov->desc.ld_tgt_count = index + 1; + cfs_waitq_init(&tgt->ltd_started); + mutex_up(&lov->lov_lock); CDEBUG(D_CONFIG, "idx=%d ltd_gen=%d ld_tgt_count=%d\n", diff --git a/lustre/osc/osc_request.c b/lustre/osc/osc_request.c index 7dd8667..cfc6ccf 100644 --- a/lustre/osc/osc_request.c +++ b/lustre/osc/osc_request.c @@ -4398,6 +4398,7 @@ static int osc_import_event(struct obd_device *obd, cli->cl_lost_grant = 0; client_obd_list_unlock(&cli->cl_loi_list_lock); ptlrpc_import_setasync(imp, -1); + obd_notify_observer(obd, obd, OBD_NOTIFY_DISCON, NULL); break; } From jeremy.filizetti at gmail.com Fri Mar 4 06:12:47 2011 From: jeremy.filizetti at gmail.com (Jeremy Filizetti) Date: Fri, 04 Mar 2011 01:12:47 -0500 Subject: [Lustre-devel] lustre 1.8+ issues with automounter In-Reply-To: References: <4D706F39.2070807@gmail.com> Message-ID: <4D7082DF.2040705@gmail.com> An example is below with some comments and a handful of the log removed. I don't actually have this many OSTs but I just created a lot of OSTs to easily reproduce the problem in a VM. autofs is setup to mount lustre. The autofs attempts to mount the file system when I typed "ls -l /lustre/xen1/tmp/testfile" where testfile is allocated on the 192nd OST IIRC. Mount kicked off by the above command by the automounter. 00000020:01200004:2:1298954011.295906:0:8398:0:(obd_mount.c:2001:lustre_fill_super()) VFS Op: sb ffff8801e7e22c00 00000020:01000004:2:1298954011.295920:0:8398:0:(obd_mount.c:2015:lustre_fill_super()) Mounting client xen1-client 00000080:00200000:2:1298954011.301889:0:8398:0:(llite_lib.c:1017:ll_fill_super()) VFS Op: sb ffff8801e7e22c00 00000080:01000000:2:1298954011.431273:0:8398:0:(llite_lib.c:1115:ll_fill_super()) Found profile xen1-client: mdc=xen1-MDT0000-mdc osc=xen1-clilov 00000080:00000010:2:1298954011.431274:0:8398:0:(llite_lib.c:1118:ll_fill_super()) kmalloced 'osc': 29 at ffff8801e7efd9a0. 00000080:00000010:2:1298954011.431276:0:8398:0:(llite_lib.c:1124:ll_fill_super()) kmalloced 'mdc': 34 at ffff8801dcb56ec0. 00000080:00000010:2:1298954011.431277:0:8398:0:(llite_lib.c:267:client_common_fill_super()) kmalloced 'data': 72 at ffff8801e9deedc0. 00000080:00100000:2:1298954011.432116:0:8398:0:(llite_lib.c:409:client_common_fill_super()) ocd_connect_flags: 0xe1440478 ocd_version: 17302784 ocd_grant: 0 00020000:01000000:1:1298954011.432928:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST0000_UUID active 00020000:01000000:1:1298954011.432977:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST0002_UUID active 00020000:01000000:1:1298954011.433025:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST0004_UUID active . . . 00020000:01000000:2:1298954011.455806:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST0094_UUID active 00020000:01000000:2:1298954011.455924:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST0095_UUID active 00020000:01000000:2:1298954011.456042:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST0096_UUID active 00020000:01000000:2:1298954011.456161:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST0097_UUID active 00020000:01000000:2:1298954011.457417:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST0098_UUID active 00000080:00000004:1:1298954011.457543:0:8398:0:(llite_lib.c:467:client_common_fill_super()) rootfid 16:[0x10:0xababf859:0x4000] 00020000:01000000:2:1298954011.457573:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST0099_UUID active 00020000:01000000:2:1298954011.457705:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST009a_UUID active 00000080:00000010:1:1298954011.457830:0:8398:0:(super25.c:57:ll_alloc_inode()) slab-alloced '(lli)': 928 at ffff8801e0de4bc0. 00020000:01000000:2:1298954011.457855:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST009b_UUID active 00000080:00000010:1:1298954011.457938:0:8398:0:(llite_lib.c:528:client_common_fill_super()) kfreed 'data': 72 at ffff8801e9deedc0. 00000080:00000010:1:1298954011.457977:0:8398:0:(llite_lib.c:1151:ll_fill_super()) kfreed 'mdc': 34 at ffff8801dcb56ec0. 00000080:00000010:1:1298954011.457979:0:8398:0:(llite_lib.c:1153:ll_fill_super()) kfreed 'osc': 29 at ffff8801e7efd9a0. 00000080:02000400:1:1298954011.457979:0:8398:0:(llite_lib.c:1157:ll_fill_super()) Client xen1-client has started 00000020:00000004:1:1298954011.457980:0:8398:0:(obd_mount.c:2053:lustre_fill_super()) Mount 192.168.66.2 at tcp8:/xen1 complete We just returned from filling the super block so now the file system is accessible, but as you can see by the lov_set_osc_active not all OSC's have been set active yet. 00020000:01000000:2:1298954011.457981:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST009c_UUID active 00020000:01000000:2:1298954011.458108:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST009d_UUID active . . . 00020000:01000000:2:1298954011.460053:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST00ac_UUID active 00020000:01000000:2:1298954011.460187:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST00ad_UUID active 00000080:00000010:1:1298954011.461272:0:8395:0:(super25.c:57:ll_alloc_inode()) slab-alloced '(lli)': 928 at ffff8801e0de4800. 00020000:01000000:2:1298954011.461487:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST00ae_UUID active 00000080:00000010:1:1298954011.461589:0:8395:0:(super25.c:57:ll_alloc_inode()) slab-alloced '(lli)': 928 at ffff8801e0de4440. 00000080:00010000:1:1298954011.461624:0:8395:0:(file.c:965:ll_glimpse_size()) Glimpsing inode 218 00000080:00020000:1:1298954011.461636:0:8395:0:(file.c:995:ll_glimpse_size()) obd_enqueue returned rc -5, returning -EIO Now glimpsing the inode from above that is allocated on xen-OST00bf which is not yet active so the set is empty and returns -EIO. 00020000:01000000:2:1298954011.461644:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST00af_UUID active 00020000:01000000:2:1298954011.461782:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST00b0_UUID active . . . 00020000:01000000:2:1298954011.463766:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST00be_UUID active 00020000:01000000:2:1298954011.463911:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) Marking OSC xen1-OST00bf_UUID active Finally the last OSC is set active, this is where client_common_fill_super should, ll_fill_super, lustre_fill_super should return from the mount syscall because the file system is now all accessible. I will take a look at your suggestion below tomorrow to see if it will handle this situate. Thanks, Jeremy > you patch is wrong in case some OSC targets will be inaccessible (in maintenance, or network troubles). > In that case lov_connect will stick in waiting for infinity time, but that is don't expected behavior. > Can you provide more details about what is situation confuses automount ? > or try to move >>> > err = obd_statfs(obd, &osfs, cfs_time_current_64() - HZ, 0); > if (err) > GOTO(out_mdc, err); >>> > from current location to something after get root fid. > > if FS mounted without lazystatfs option, obd_statfs will blocked until all connection requests is finished. > so you will have same behavior but without changes in obd_connect() code. From alexey.lyashkov at clusterstor.com Fri Mar 4 06:21:38 2011 From: alexey.lyashkov at clusterstor.com (Alexey Lyashkov) Date: Fri, 4 Mar 2011 09:21:38 +0300 Subject: [Lustre-devel] lustre 1.8+ issues with automounter In-Reply-To: <4D7082DF.2040705@gmail.com> References: <4D706F39.2070807@gmail.com> <4D7082DF.2040705@gmail.com> Message-ID: <6BA20898-9773-4E43-A697-C31C0A2C00E4@clusterstor.com> if you can add "df " call after mounting lustre fs - it will also help. On Mar 4, 2011, at 09:12, Jeremy Filizetti wrote: > An example is below with some comments and a handful of the log > removed. I don't actually have this many OSTs but I just created a lot > of OSTs to easily reproduce the problem in a VM. autofs is setup to > mount lustre. The autofs attempts to mount the file system when I typed > "ls -l /lustre/xen1/tmp/testfile" where testfile is allocated on the > 192nd OST IIRC. > > Mount kicked off by the above command by the automounter. > 00000020:01200004:2:1298954011.295906:0:8398:0:(obd_mount.c:2001:lustre_fill_super()) > VFS Op: sb ffff8801e7e22c00 > 00000020:01000004:2:1298954011.295920:0:8398:0:(obd_mount.c:2015:lustre_fill_super()) > Mounting client xen1-client > 00000080:00200000:2:1298954011.301889:0:8398:0:(llite_lib.c:1017:ll_fill_super()) > VFS Op: sb ffff8801e7e22c00 > 00000080:01000000:2:1298954011.431273:0:8398:0:(llite_lib.c:1115:ll_fill_super()) > Found profile xen1-client: mdc=xen1-MDT0000-mdc osc=xen1-clilov > 00000080:00000010:2:1298954011.431274:0:8398:0:(llite_lib.c:1118:ll_fill_super()) > kmalloced 'osc': 29 at ffff8801e7efd9a0. > 00000080:00000010:2:1298954011.431276:0:8398:0:(llite_lib.c:1124:ll_fill_super()) > kmalloced 'mdc': 34 at ffff8801dcb56ec0. > 00000080:00000010:2:1298954011.431277:0:8398:0:(llite_lib.c:267:client_common_fill_super()) > kmalloced 'data': 72 at ffff8801e9deedc0. > 00000080:00100000:2:1298954011.432116:0:8398:0:(llite_lib.c:409:client_common_fill_super()) > ocd_connect_flags: 0xe1440478 ocd_version: 17302784 ocd_grant: 0 > 00020000:01000000:1:1298954011.432928:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST0000_UUID active > 00020000:01000000:1:1298954011.432977:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST0002_UUID active > 00020000:01000000:1:1298954011.433025:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST0004_UUID active > . > . > . > 00020000:01000000:2:1298954011.455806:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST0094_UUID active > 00020000:01000000:2:1298954011.455924:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST0095_UUID active > 00020000:01000000:2:1298954011.456042:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST0096_UUID active > 00020000:01000000:2:1298954011.456161:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST0097_UUID active > 00020000:01000000:2:1298954011.457417:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST0098_UUID active > 00000080:00000004:1:1298954011.457543:0:8398:0:(llite_lib.c:467:client_common_fill_super()) > rootfid 16:[0x10:0xababf859:0x4000] > 00020000:01000000:2:1298954011.457573:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST0099_UUID active > 00020000:01000000:2:1298954011.457705:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST009a_UUID active > 00000080:00000010:1:1298954011.457830:0:8398:0:(super25.c:57:ll_alloc_inode()) > slab-alloced '(lli)': 928 at ffff8801e0de4bc0. > 00020000:01000000:2:1298954011.457855:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST009b_UUID active > 00000080:00000010:1:1298954011.457938:0:8398:0:(llite_lib.c:528:client_common_fill_super()) > kfreed 'data': 72 at ffff8801e9deedc0. > 00000080:00000010:1:1298954011.457977:0:8398:0:(llite_lib.c:1151:ll_fill_super()) > kfreed 'mdc': 34 at ffff8801dcb56ec0. > 00000080:00000010:1:1298954011.457979:0:8398:0:(llite_lib.c:1153:ll_fill_super()) > kfreed 'osc': 29 at ffff8801e7efd9a0. > 00000080:02000400:1:1298954011.457979:0:8398:0:(llite_lib.c:1157:ll_fill_super()) > Client xen1-client has started > 00000020:00000004:1:1298954011.457980:0:8398:0:(obd_mount.c:2053:lustre_fill_super()) > Mount 192.168.66.2 at tcp8:/xen1 complete > > We just returned from filling the super block so now the file system is > accessible, but as you can see by the lov_set_osc_active not all OSC's > have been set active yet. > > 00020000:01000000:2:1298954011.457981:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST009c_UUID active > 00020000:01000000:2:1298954011.458108:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST009d_UUID active > . > . > . > 00020000:01000000:2:1298954011.460053:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST00ac_UUID active > 00020000:01000000:2:1298954011.460187:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST00ad_UUID active > 00000080:00000010:1:1298954011.461272:0:8395:0:(super25.c:57:ll_alloc_inode()) > slab-alloced '(lli)': 928 at ffff8801e0de4800. > 00020000:01000000:2:1298954011.461487:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST00ae_UUID active > 00000080:00000010:1:1298954011.461589:0:8395:0:(super25.c:57:ll_alloc_inode()) > slab-alloced '(lli)': 928 at ffff8801e0de4440. > 00000080:00010000:1:1298954011.461624:0:8395:0:(file.c:965:ll_glimpse_size()) > Glimpsing inode 218 > 00000080:00020000:1:1298954011.461636:0:8395:0:(file.c:995:ll_glimpse_size()) > obd_enqueue returned rc -5, returning -EIO > > Now glimpsing the inode from above that is allocated on xen-OST00bf > which is not yet active so the set is empty and returns -EIO. > > 00020000:01000000:2:1298954011.461644:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST00af_UUID active > 00020000:01000000:2:1298954011.461782:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST00b0_UUID active > . > . > . > 00020000:01000000:2:1298954011.463766:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST00be_UUID active > 00020000:01000000:2:1298954011.463911:0:11545:0:(lov_obd.c:570:lov_set_osc_active()) > Marking OSC xen1-OST00bf_UUID active > > Finally the last OSC is set active, this is where > client_common_fill_super should, ll_fill_super, lustre_fill_super should > return from the mount syscall because the file system is now all accessible. > > I will take a look at your suggestion below tomorrow to see if it will > handle this situate. > > > Thanks, > Jeremy > >> you patch is wrong in case some OSC targets will be inaccessible (in maintenance, or network troubles). >> In that case lov_connect will stick in waiting for infinity time, but that is don't expected behavior. >> Can you provide more details about what is situation confuses automount ? >> or try to move >>>> >> err = obd_statfs(obd, &osfs, cfs_time_current_64() - HZ, 0); >> if (err) >> GOTO(out_mdc, err); >>>> >> from current location to something after get root fid. >> >> if FS mounted without lazystatfs option, obd_statfs will blocked until all connection requests is finished. >> so you will have same behavior but without changes in obd_connect() code. > ______________________________________________________________________ This email may contain privileged or confidential information, which should only be used for the purpose for which it was sent by Xyratex. No further rights or licenses are granted to use such information. If you are not the intended recipient of this message, please notify the sender by return and delete it. You may not use, copy, disclose or rely on the information contained in it. Internet email is susceptible to data corruption, interception and unauthorised amendment for which Xyratex does not accept liability. While we have taken reasonable precautions to ensure that this email is free of viruses, Xyratex does not accept liability for the presence of any computer viruses in this email, nor for any losses caused as a result of viruses. Xyratex Technology Limited (03134912), Registered in England & Wales, Registered Office, Langstone Road, Havant, Hampshire, PO9 1SA. The Xyratex group of companies also includes, Xyratex Ltd, registered in Bermuda, Xyratex International Inc, registered in California, Xyratex (Malaysia) Sdn Bhd registered in Malaysia, Xyratex Technology (Wuxi) Co Ltd registered in The People's Republic of China and Xyratex Japan Limited registered in Japan. ______________________________________________________________________ From adilger at whamcloud.com Fri Mar 4 06:39:30 2011 From: adilger at whamcloud.com (Andreas Dilger) Date: Thu, 3 Mar 2011 23:39:30 -0700 Subject: [Lustre-devel] lustre 1.8+ issues with automounter In-Reply-To: <4D706F39.2070807@gmail.com> References: <4D706F39.2070807@gmail.com> Message-ID: <9F9F3E45-4B9C-42F4-9D7F-000C70E4BAC9@whamcloud.com> On 2011-03-03, at 9:48 PM, Jeremy Filizetti wrote: > Ever since we moved from Lustre 1.6.6 to 1.8 I've seen issues with using > the automounter and Lustre. I've finally got around to looking at what > the issue is, but I'm not quite sure what the correct way to resolve it > is. I think the issue will remain in 2.0+ but I didn't look closely at > the code. Interesting. I've known about automount problems with Lustre for some time (probably a search in the list history would find a bunch), but nobody has every dug into the root cause. Thanks for taking the time to investigate. > The issue is that lov_connect which calls lov_connect_obd is > an asynchronous connect that does not wait for all OSCs to be connected > before returning. In the end lustre_fill_super can return before all > OSCs have been set active so any file operations that caused the > automount may return an error. Many lov functions check to make sure > the lov_tgt_desc ltd_active flag is 1 or return -EIO. Right. This is to allow Lustre to operate in "failout" mode (i.e. never wait for recovery on a down OST, and instead allow the application to do something else), and/or if the administrator marks the OST unavailable via "lctl deactivate" if it is down for some extended period (major hardware failure, corruption, etc). > The following patch handles things correctly by waiting until all OSC's > that are set to be activated are active before returning from filling > the super block. There are a few problems that I'm not sure of what the > expected results are with Lustre. For example if an OST has not been > mounted the client will attempt to connect and end up returning -ENODEV > and setting the import_state as LUSTRE_IMP_DISCON. Without the patch > the client mounts immediately even though the OSC is unavailable, with > it the mount would not return until the user kills the process, the OBD > is set inactive, or the state changes. This is done intentionally, so that the client can complete the mount without waiting for all of the connections, which may take tens of seconds when there are 100k of clients booting at the same time, or may take a very long time if the OST is down, and block the client boot process indefinitely. > To provide the same functionality an extra condition would need to be added > to the l_wait_event condition to monitor the import state is not connecting. > However if I do that, I'm not sure things handle failover nodes correctly. > So what I'm wondering is what are the expected actions for the different > conditions of OSTs. I wonder if it makes sense to start the OSCs in "active" mode, and only mark them inactive if they fail the initial connect request. I haven't looked at this code for a long time, so I'm not sure if this will have some unintended side effects. For future patch submissions, please follow the Lustre Coding Guidelines at http://wiki.lustre.org/index.php/Coding_Guidelines > diff --git a/lustre/include/obd.h b/lustre/include/obd.h > index e89805d..3046a5c 100644 > --- a/lustre/include/obd.h > +++ b/lustre/include/obd.h > @@ -754,6 +754,8 @@ struct lov_tgt_desc { > unsigned long ltd_active:1,/* is this target up for > requests */ > ltd_activate:1,/* should this target be > activated */ > ltd_reap:1; /* should this target be > deleted */ > + cfs_waitq_t ltd_started; /* waitqueue to notify tgt has > been fully started > + * so IO can start */ > }; > > /* Pool metadata */ > @@ -942,6 +944,8 @@ enum obd_notify_event { > OBD_NOTIFY_ACTIVE, > /* Device deactivated */ > OBD_NOTIFY_INACTIVE, > + /* Device disconnected */ > + OBD_NOTIFY_DISCON, > /* Connect data for import were changed */ > OBD_NOTIFY_OCD, > /* Sync request */ > diff --git a/lustre/lov/lov_obd.c b/lustre/lov/lov_obd.c > index 8b2d848..ff4a04a 100644 > --- a/lustre/lov/lov_obd.c > +++ b/lustre/lov/lov_obd.c > @@ -222,7 +222,33 @@ static int lov_notify(struct obd_device *obd, > struct obd_device *watched, > } > /* active event should be pass lov target index as data */ > data = &rc; > - } > + } else if (ev == OBD_NOTIFY_DISCON) { > + struct lov_tgt_desc *tgt; > + struct lov_obd *lov = &obd->u.lov; > + int i; > + > + LASSERT(watched); > + if (strcmp(watched->obd_type->typ_name, LUSTRE_OSC_NAME)) { > + CERROR("unexpected notification of %s %s!\n", > + watched->obd_type->typ_name, > + watched->obd_name); > + RETURN(-EINVAL); > + } > + > + obd_getref(obd); > + for (i = 0; i < lov->desc.ld_tgt_count; i++) { > + tgt = lov->lov_tgts[i]; > + if (!tgt || !tgt->ltd_exp) > + continue; > + > + if (obd_uuid_equals(&watched->u.cli.cl_target_uuid, > &tgt->ltd_uuid)) { > + cfs_waitq_signal(&lov->lov_tgts[i]->ltd_started); > + data = &i; > + break; > + } > + } > + obd_putref(obd); > + } > > /* Pass the notification up the chain. */ > if (watched) { > @@ -424,6 +450,27 @@ static int lov_connect(struct lustre_handle *conn, > struct obd_device *obd, > obd->obd_name, rc); > } > } > + > + /* Wait for all the connections to complete before returning so > that all > + * obds are set active that should be. Otherwise IO that > happens immediately > + * after mount could (autofs) could glimpse or touch objects before > the connecction > + * is established */ > + for (i = 0; i < lov->desc.ld_tgt_count; i++) { > + struct l_wait_info lwi = { 0 }; > + > + tgt = lov->lov_tgts[i]; > + if (!tgt || !tgt->ltd_exp || obd_uuid_empty(&tgt->ltd_uuid)) > + continue; > + > + if (tgt->ltd_activate == tgt->ltd_active) > + continue; > + > + CDEBUG(D_CONFIG, "Target %s activate/active %d/%d, waiting on > state change\n", > + tgt->ltd_obd->obd_name, tgt->ltd_activate, tgt->ltd_active); > + > + l_wait_event(tgt->ltd_started, tgt->ltd_activate == > tgt->ltd_active || > + tgt->ltd_obd->u.cli.cl_import->imp_deactive, &lwi); > + } > obd_putref(obd); > > RETURN(0); > @@ -445,6 +492,9 @@ static int lov_disconnect_obd(struct obd_device > *obd, struct lov_tgt_desc *tgt) > tgt->ltd_active = 0; > lov->desc.ld_active_tgt_count--; > tgt->ltd_exp->exp_obd->obd_inactive = 1; > + > + /* If state change wake up wait queue */ > + cfs_waitq_signal(&tgt->ltd_started); > } > > lov_proc_dir = lprocfs_srch(obd->obd_proc_entry, "target_obds"); > @@ -582,6 +632,9 @@ static int lov_set_osc_active(struct obd_device > *obd, struct obd_uuid *uuid, > lov->lov_tgts[i]->ltd_qos.ltq_penalty = 0; > > out: > + if (i >= 0) > + cfs_waitq_signal(&lov->lov_tgts[i]->ltd_started); > + > obd_putref(obd); > RETURN(i); > } > @@ -673,6 +726,8 @@ static int lov_add_target(struct obd_device *obd, > struct obd_uuid *uuidp, > if (index >= lov->desc.ld_tgt_count) > lov->desc.ld_tgt_count = index + 1; > > + cfs_waitq_init(&tgt->ltd_started); > + > mutex_up(&lov->lov_lock); > > CDEBUG(D_CONFIG, "idx=%d ltd_gen=%d ld_tgt_count=%d\n", > diff --git a/lustre/osc/osc_request.c b/lustre/osc/osc_request.c > index 7dd8667..cfc6ccf 100644 > --- a/lustre/osc/osc_request.c > +++ b/lustre/osc/osc_request.c > @@ -4398,6 +4398,7 @@ static int osc_import_event(struct obd_device *obd, > cli->cl_lost_grant = 0; > client_obd_list_unlock(&cli->cl_loi_list_lock); > ptlrpc_import_setasync(imp, -1); > + obd_notify_observer(obd, obd, OBD_NOTIFY_DISCON, NULL); > > break; > } > > _______________________________________________ > Lustre-devel mailing list > Lustre-devel at lists.lustre.org > http://lists.lustre.org/mailman/listinfo/lustre-devel Cheers, Andreas -- Andreas Dilger Principal Engineer Whamcloud, Inc. From alexey.lyashkov at clusterstor.com Fri Mar 4 09:22:14 2011 From: alexey.lyashkov at clusterstor.com (Alexey Lyashkov) Date: Fri, 4 Mar 2011 12:22:14 +0300 Subject: [Lustre-devel] lustre 1.8+ issues with automounter In-Reply-To: <9F9F3E45-4B9C-42F4-9D7F-000C70E4BAC9@whamcloud.com> References: <4D706F39.2070807@gmail.com> <9F9F3E45-4B9C-42F4-9D7F-000C70E4BAC9@whamcloud.com> Message-ID: On Mar 4, 2011, at 09:39, Andreas Dilger wrote: > On 2011-03-03, at 9:48 PM, Jeremy Filizetti wrote: >> Ever since we moved from Lustre 1.6.6 to 1.8 I've seen issues with using >> the automounter and Lustre. I've finally got around to looking at what >> the issue is, but I'm not quite sure what the correct way to resolve it >> is. I think the issue will remain in 2.0+ but I didn't look closely at >> the code. > > Interesting. I've known about automount problems with Lustre for some time (probably a search in the list history would find a bunch), but nobody has every dug into the root cause. Thanks for taking the time to investigate. > Looks it is result of rq_no_resend flag for glimpse request, so it will failed (instead of put to delay list) and that error returned to caller. -------------------------------------- Alexey Lyashkov alexey.lyashkov at clusterstor.com ______________________________________________________________________ This email may contain privileged or confidential information, which should only be used for the purpose for which it was sent by Xyratex. No further rights or licenses are granted to use such information. If you are not the intended recipient of this message, please notify the sender by return and delete it. You may not use, copy, disclose or rely on the information contained in it. Internet email is susceptible to data corruption, interception and unauthorised amendment for which Xyratex does not accept liability. While we have taken reasonable precautions to ensure that this email is free of viruses, Xyratex does not accept liability for the presence of any computer viruses in this email, nor for any losses caused as a result of viruses. Xyratex Technology Limited (03134912), Registered in England & Wales, Registered Office, Langstone Road, Havant, Hampshire, PO9 1SA. The Xyratex group of companies also includes, Xyratex Ltd, registered in Bermuda, Xyratex International Inc, registered in California, Xyratex (Malaysia) Sdn Bhd registered in Malaysia, Xyratex Technology (Wuxi) Co Ltd registered in The People's Republic of China and Xyratex Japan Limited registered in Japan. ______________________________________________________________________ From alexey_lyashkov at xyratex.com Fri Mar 4 05:47:59 2011 From: alexey_lyashkov at xyratex.com (Alexey Lyashkov) Date: Fri, 4 Mar 2011 08:47:59 +0300 Subject: [Lustre-devel] lustre 1.8+ issues with automounter In-Reply-To: <4D706F39.2070807@gmail.com> References: <4D706F39.2070807@gmail.com> Message-ID: On Mar 4, 2011, at 07:48, Jeremy Filizetti wrote: > Ever since we moved from Lustre 1.6.6 to 1.8 I've seen issues with using > the automounter and Lustre. I've finally got around to looking at what > the issue is, but I'm not quite sure what the correct way to resolve it > is. I think the issue will remain in 2.0+ but I didn't look closely at > the code. The issue is that lov_connect which calls lov_connect_obd is > an asynchronous connect that does not wait for all OSCs to be connected > before returning. In the end lustre_fill_super can return before all > OSCs have been set active so any file operations that caused the > automount may return an error. Many lov functions check to make sure > the lov_tgt_desc ltd_active flag is 1 or return -EIO. > > you patch is wrong in case some OSC targets will be inaccessible (in maintenance, or network troubles). In that case lov_connect will stick in waiting for infinity time, but that is don't expected behavior. Can you provide more details about what is situation confuses automount ? or try to move >> err = obd_statfs(obd, &osfs, cfs_time_current_64() - HZ, 0); if (err) GOTO(out_mdc, err); >> from current location to something after get root fid. if FS mounted without lazystatfs option, obd_statfs will blocked until all connection requests is finished. so you will have same behavior but without changes in obd_connect() code. -------------------------------------------- Alexey Lyashkov alexey_lyashkov at xyratex.com ______________________________________________________________________ This email may contain privileged or confidential information, which should only be used for the purpose for which it was sent by Xyratex. No further rights or licenses are granted to use such information. If you are not the intended recipient of this message, please notify the sender by return and delete it. You may not use, copy, disclose or rely on the information contained in it. Internet email is susceptible to data corruption, interception and unauthorised amendment for which Xyratex does not accept liability. While we have taken reasonable precautions to ensure that this email is free of viruses, Xyratex does not accept liability for the presence of any computer viruses in this email, nor for any losses caused as a result of viruses. Xyratex Technology Limited (03134912), Registered in England & Wales, Registered Office, Langstone Road, Havant, Hampshire, PO9 1SA. The Xyratex group of companies also includes, Xyratex Ltd, registered in Bermuda, Xyratex International Inc, registered in California, Xyratex (Malaysia) Sdn Bhd registered in Malaysia, Xyratex Technology (Wuxi) Co Ltd registered in The People's Republic of China and Xyratex Japan Limited registered in Japan. ______________________________________________________________________ From raviprakashdrbh at aol.com Sat Mar 5 01:30:26 2011 From: raviprakashdrbh at aol.com (Ravi) Date: Fri, 4 Mar 2011 20:30:26 -0500 (EST) Subject: [Lustre-devel] LNetPoll undefined Message-ID: <8CDA8EC0A0C0017-A90-CEFA@webmail-d097.sysops.aol.com> Hello I am using LNetWait (blocking call ) on a particular event .After i recevie this event i break from the loop which waits for this event and proceed but when another event is added into the event queue the system crashes.I thought LNetPoll would be better as i can just poll for that particular event without disturbing the event queue but when i make i get undefined.Any thoughts . Thanks Ravi -------------- next part -------------- An HTML attachment was scrubbed... URL: From raviprakashdrbh at aol.com Sat Mar 5 01:37:38 2011 From: raviprakashdrbh at aol.com (Ravi) Date: Fri, 4 Mar 2011 20:37:38 -0500 (EST) Subject: [Lustre-devel] LnetEQPoll undefined Message-ID: <8CDA8ED0BAAFC11-A90-CFB2@webmail-d097.sysops.aol.com> Hello The error before crashing is Mar 4 20:35:27 ws11 kernel: BUG: soft lockup - CPU#0 stuck for 10s! [insmod:4609] searched a lot for what it means but couldnt figure out. Thanks -------------- next part -------------- An HTML attachment was scrubbed... URL: From liang at whamcloud.com Sat Mar 5 01:54:09 2011 From: liang at whamcloud.com (Liang Zhen) Date: Sat, 5 Mar 2011 09:54:09 +0800 Subject: [Lustre-devel] LNetPoll undefined In-Reply-To: <8CDA8EC0A0C0017-A90-CEFA@webmail-d097.sysops.aol.com> References: <8CDA8EC0A0C0017-A90-CEFA@webmail-d097.sysops.aol.com> Message-ID: <7E84F671-5637-4253-95BB-6999A17F3C4E@whamcloud.com> Hi Ravi, Which version of Lustre/LNet are you trying with? Are you trying to build some new code over LNet? Could you show us some example code if you don't mind? btw, If you are trying this in kernel space, I would suggest to use eq_callback (LNetEQAlloc(...eq_callback)) instead of LNetEQPoll/LNetEQWait, which is better for performance. Polling is not good for performance because all EQs share one single waitq in LNet. Regards Liang On Mar 5, 2011, at 9:30 AM, Ravi wrote: > Hello > > I am using LNetWait (blocking call ) on a particular event .After i recevie this event i break from the loop which waits for this event and proceed but when another event is added into the event queue the system crashes.I thought LNetPoll would be better as i can just poll for that particular event without disturbing the event queue but when i make i get undefined.Any thoughts . > > Thanks > Ravi > > _______________________________________________ > Lustre-devel mailing list > Lustre-devel at lists.lustre.org > http://lists.lustre.org/mailman/listinfo/lustre-devel -------------- next part -------------- An HTML attachment was scrubbed... URL: From raviprakashdrbh at aol.com Sat Mar 5 18:51:01 2011 From: raviprakashdrbh at aol.com (Ravi) Date: Sat, 5 Mar 2011 13:51:01 -0500 (EST) Subject: [Lustre-devel] LNetPoll undefined In-Reply-To: <7E84F671-5637-4253-95BB-6999A17F3C4E@whamcloud.com> References: <8CDA8EC0A0C0017-A90-CEFA@webmail-d097.sysops.aol.com> <7E84F671-5637-4253-95BB-6999A17F3C4E@whamcloud.com> Message-ID: <8CDA97D68228C90-21C-11D97@webmail-m099.sysops.aol.com> Hello Thanks for the reply.I am using lustre-1.8.1.1 .I am working on a new module which is used for some kind of delegation operations.I am using Lnet operations in this module. The code which fails is do { rc = LNetEQWait(lnet_eq_hd, &ev); if( ev.type == LNET_EVENT_PUT ) break; } while ( rc != 0); Here i am waiting on some PUT event from a client and then break from the loop.And do some operations accordingly.But next time when i perform some PUT operation (for example) and it gets logged into the event queue i try reading from that event but the MDS fails. I also tried using these functions //rc = LNetEQPoll( &lnet_eq_hd, 1,2000, &ev, &which); in place of LNetEQWait but it says undefined . Can you please throw some light on eq_callback function as i havnt found it in Lnet manual to go through. The log before crashing shows : Mar 4 20:35:08 ws11 kernel: type=LNET_EVENT_SEND, pt-idx=53, mbits=0x1234abcd, rlen=64, mlen=64, md.user_ptr=0xaaaabbbb, hdr-data=0x0 Mar 4 20:35:08 ws11 kernel: status=0, unlnk=0, offset=0, seq=2 //Iam asssuming it fails here as till here it prints fine.I also want to mention that this operation is successful as well. Mar 4 20:35:27 ws11 kernel: BUG: soft lockup - CPU#0 stuck for 10s! [insmod:4609] Mar 4 20:35:27 ws11 kernel: CPU 0: Mar 4 20:35:27 ws11 kernel: Modules linked in: tmod(U) ksocklnd(U) ko2iblnd(FU) lnet(U) libcfs(U) autofs4(U) hidp(U) nfs(U) fscache(U) nfs_acl(U) rfcomm(U) l2cap(U) bluetooth(U) lockd(U) sunrpc(U) cpufreq_ondemand(U) acpi_cpufreq(U) freq_table(U) ip_conntrack_netbios_ns(U) ipt_REJECT(U) xt_state(U) ip_conntrack(U) nfnetlink(U) iptable_filter(U) ip_tables(U) ip6t_REJECT(U) xt_tcpudp(U) ip6table_filter(U) ip6_tables(U) x_tables(U) rdma_ucm(U) ib_sdp(U) rdma_cm(U) iw_cm(U) ib_addr(U) ib_ipoib(U) ipoib_helper(U) ib_cm(U) ib_sa(U) ipv6(U) xfrm_nalgo(U) crypto_api(U) ib_uverbs(U) ib_umad(U) mlx4_en(U) mlx4_ib(U) mlx4_core(U) loop(U) dm_multipath(U) scsi_dh(U) video(U) hwmon(U) backlight(U) sbs(U) i2c_ec(U) button(U) battery(U) asus_acpi(U) acpi_memhotplug(U) ac(U) lp(U) snd_hda_intel(U) snd_seq_dummy(U) snd_seq_oss(U) snd_seq_midi_event(U) snd_seq(U) snd_seq_device(U) snd_pcm_oss(U) snd_mixer_oss(U) ib_mthca(U) snd_pcm(U) snd_timer(U) snd_page_alloc(U) ib_mad(U) snd_hwdep(U) snd(U) sg(U) ib_core(U) e100(U) ide_cd( Mar 4 20:35:27 ws11 kernel: ) mii(U) serio_raw(U) pcspkr(U) i2c_i801(U) cdrom(U) soundcore(U) parport_pc(U) shpchp(U) i2c_core(U) parport(U) dm_raid45(U) dm_message(U) dm_region_hash(U) dm_mem_cache(U) dm_snapshot(U) dm_zero(U) dm_mirror(U) dm_log(U) dm_mod(U) ata_piix(U) libata(U) sd_mod(U) scsi_mod(U) ext3(U) jbd(U) uhci_hcd(U) ohci_hcd(U) ehci_hcd(U) Mar 4 20:35:27 ws11 kernel: Pid: 4609, comm: insmod Tainted: GF 2.6.18-128.7.1.el5-lustre.1.8.1.1smp-cust #2 Mar 4 20:35:27 ws11 kernel: RIP: 0010:[] [] .text.lock.spinlock+0x2/0x30 Mar 4 20:35:27 ws11 kernel: RSP: 0018:ffff810021099d10 EFLAGS: 00000286 Mar 4 20:35:27 ws11 kernel: RAX: 0000000000000002 RBX: 00000000ffffffff RCX: ffff810021099df8 Mar 4 20:35:27 ws11 kernel: RDX: 0000000000000001 RSI: 0000000000000001 RDI: ffffffff8888a7a0 Mar 4 20:35:27 ws11 kernel: RBP: ffff81003a04e1c0 R08: ffff810021099ddc R09: 000000001234abcd ....... I hope this helps Thanks -----Original Message----- From: Liang Zhen To: Ravi Cc: lustre-devel Sent: Fri, Mar 4, 2011 8:54 pm Subject: Re: [Lustre-devel] LNetPoll undefined Hi Ravi, Which version of Lustre/LNet are you trying with? Are you trying to build some new code over LNet? Could you show us some example code if you don't mind? btw, If you are trying this in kernel space, I would suggest to use eq_callback (LNetEQAlloc(...eq_callback)) instead of LNetEQPoll/LNetEQWait, which is better for performance. Polling is not good for performance because all EQs share one single waitq in LNet. Regards Liang On Mar 5, 2011, at 9:30 AM, Ravi wrote: Hello I am using LNetWait (blocking call ) on a particular event .After i recevie this event i break from the loop which waits for this event and proceed but when another event is added into the event queue the system crashes.I thought LNetPoll would be better as i can just poll for that particular event without disturbing the event queue but when i make i get undefined.Any thoughts . Thanks Ravi _______________________________________________ Lustre-devel mailing list Lustre-devel at lists.lustre.org http://lists.lustre.org/mailman/listinfo/lustre-devel -------------- next part -------------- An HTML attachment was scrubbed... URL: From liang at whamcloud.com Mon Mar 7 01:35:29 2011 From: liang at whamcloud.com (Liang Zhen) Date: Mon, 7 Mar 2011 09:35:29 +0800 Subject: [Lustre-devel] LNetPoll undefined In-Reply-To: <8CDA97D68228C90-21C-11D97@webmail-m099.sysops.aol.com> References: <8CDA8EC0A0C0017-A90-CEFA@webmail-d097.sysops.aol.com> <7E84F671-5637-4253-95BB-6999A17F3C4E@whamcloud.com> <8CDA97D68228C90-21C-11D97@webmail-m099.sysops.aol.com> Message-ID: Ravi, I think the soft lockup probably is because the thread is polling on the EQ and expecting EVENT_PUT, however, there are a lot of EVENT_SEND(server keeps calling LNetPut or LNetGet with MD on the same EQ?) so it became to a busy loop which is always trying to get LNET_LOCK to poll new event, and kernel can't schedule watchdog on some CPUs then raise the warning. pseudo code for eq_callback is like wait_queue_head_t my_waitq; struct list_head req_list; void my_callback(lnet_event_t *ev) { if (ev->type != LNET_EVENT_PUT) return; /* construct request form data in MD */ req = ....; ... add_req_to_queue(req, req_list); wake_up(my_waitq); } my_thread() { ... rc = LNetEQAlloc(1024, my_callback, &eqh); ... while (1) { while (!list_empty(&req_list)) { req = list_entry(req_list.next, ...); list_del(&req->list); handle_request(req); } init_waitqueue_entry(wait, current); add_wait_queue(my_waitq, wait) if (list_empty(&req_list)) schedule(); remove_wait_queue(my_waitq, wait); } } NB: this is just pseudo code and there should be some locks to protect, if you want some real code that is using eq_callback, please lookup into lnet/selftest/rpc.c Yes, LNetEQPoll is not exported... so if you really want to use it, please just add this line to lnet/lnet/module.c: EXPORT_SYMBOL(LNetEQPoll); Though they should be exact samely for your case. Regards Liang On Mar 6, 2011, at 2:51 AM, Ravi wrote: > Hello > Thanks for the reply.I am using lustre-1.8.1.1 .I am working on a new module which is used for some kind of delegation operations.I am using Lnet operations in this module. > > > The code which fails is > do { > rc = LNetEQWait(lnet_eq_hd, &ev); > if( ev.type == LNET_EVENT_PUT ) > break; > > } while ( rc != 0); > > > > > Here i am waiting on some PUT event from a client and then break from the loop.And do some operations accordingly.But next time when i perform some PUT operation (for example) and it gets logged into the event queue i try reading from that event but the MDS fails. > I also tried using these functions //rc = LNetEQPoll( &lnet_eq_hd, 1,2000, &ev, &which); in place of LNetEQWait but it says undefined . > > Can you please throw some light on eq_callback function as i havnt found it in Lnet manual to go through. > > > The log before crashing shows : > > > Mar 4 20:35:08 ws11 kernel: type=LNET_EVENT_SEND, pt-idx=53, mbits=0x1234abcd, rlen=64, mlen=64, md.user_ptr=0xaaaabbbb, hdr-data=0x0 > Mar 4 20:35:08 ws11 kernel: status=0, unlnk=0, offset=0, seq=2 > > //Iam asssuming it fails here as till here it prints fine.I also want to mention that this operation is successful as well. > > > Mar 4 20:35:27 ws11 kernel: BUG: soft lockup - CPU#0 stuck for 10s! [insmod:4609] > Mar 4 20:35:27 ws11 kernel: CPU 0: > Mar 4 20:35:27 ws11 kernel: Modules linked in: tmod(U) ksocklnd(U) ko2iblnd(FU) lnet(U) libcfs(U) autofs4(U) hidp(U) nfs(U) fscache(U) nfs_acl(U) rfcomm(U) l2cap(U) bluetooth(U) lockd(U) sunrpc(U) cpufreq_ondemand(U) acpi_cpufreq(U) freq_table(U) ip_conntrack_netbios_ns(U) ipt_REJECT(U) xt_state(U) ip_conntrack(U) nfnetlink(U) iptable_filter(U) ip_tables(U) ip6t_REJECT(U) xt_tcpudp(U) ip6table_filter(U) ip6_tables(U) x_tables(U) rdma_ucm(U) ib_sdp(U) rdma_cm(U) iw_cm(U) ib_addr(U) ib_ipoib(U) ipoib_helper(U) ib_cm(U) ib_sa(U) ipv6(U) xfrm_nalgo(U) crypto_api(U) ib_uverbs(U) ib_umad(U) mlx4_en(U) mlx4_ib(U) mlx4_core(U) loop(U) dm_multipath(U) scsi_dh(U) video(U) hwmon(U) backlight(U) sbs(U) i2c_ec(U) button(U) battery(U) asus_acpi(U) acpi_memhotplug(U) ac(U) lp(U) snd_hda_intel(U) snd_seq_dummy(U) snd_seq_oss(U) snd_seq_midi_event(U) snd_seq(U) snd_seq_device(U) snd_pcm_oss(U) snd_mixer_oss(U) ib_mthca(U) snd_pcm(U) snd_timer(U) snd_page_alloc(U) ib_mad(U) snd_hwdep(U) snd(U) sg(U) ib_core(U) e100(U) ide_cd( > Mar 4 20:35:27 ws11 kernel: ) mii(U) serio_raw(U) pcspkr(U) i2c_i801(U) cdrom(U) soundcore(U) parport_pc(U) shpchp(U) i2c_core(U) parport(U) dm_raid45(U) dm_message(U) dm_region_hash(U) dm_mem_cache(U) dm_snapshot(U) dm_zero(U) dm_mirror(U) dm_log(U) dm_mod(U) ata_piix(U) libata(U) sd_mod(U) scsi_mod(U) ext3(U) jbd(U) uhci_hcd(U) ohci_hcd(U) ehci_hcd(U) > Mar 4 20:35:27 ws11 kernel: Pid: 4609, comm: insmod Tainted: GF 2.6.18-128.7.1.el5-lustre.1.8.1.1smp-cust #2 > Mar 4 20:35:27 ws11 kernel: RIP: 0010:[] [] .text.lock.spinlock+0x2/0x30 > Mar 4 20:35:27 ws11 kernel: RSP: 0018:ffff810021099d10 EFLAGS: 00000286 > Mar 4 20:35:27 ws11 kernel: RAX: 0000000000000002 RBX: 00000000ffffffff RCX: ffff810021099df8 > Mar 4 20:35:27 ws11 kernel: RDX: 0000000000000001 RSI: 0000000000000001 RDI: ffffffff8888a7a0 > Mar 4 20:35:27 ws11 kernel: RBP: ffff81003a04e1c0 R08: ffff810021099ddc R09: 000000001234abcd > ....... > > > > I hope this helps > > Thanks > > > > > > -----Original Message----- > From: Liang Zhen > To: Ravi > Cc: lustre-devel > Sent: Fri, Mar 4, 2011 8:54 pm > Subject: Re: [Lustre-devel] LNetPoll undefined > > Hi Ravi, > > Which version of Lustre/LNet are you trying with? Are you trying to build some new code over LNet? Could you show us some example code if you don't mind? > btw, If you are trying this in kernel space, I would suggest to use eq_callback (LNetEQAlloc(...eq_callback)) instead of LNetEQPoll/LNetEQWait, which is better for performance. Polling is not good for performance because all EQs share one single waitq in LNet. > > Regards > Liang > > On Mar 5, 2011, at 9:30 AM, Ravi wrote: > >> Hello >> >> I am using LNetWait (blocking call ) on a particular event .After i recevie this event i break from the loop which waits for this event and proceed but when another event is added into the event queue the system crashes.I thought LNetPoll would be better as i can just poll for that particular event without disturbing the event queue but when i make i get undefined.Any thoughts . >> >> Thanks >> Ravi >> >> _______________________________________________ >> Lustre-devel mailing list >> Lustre-devel at lists.lustre.org >> http://lists.lustre.org/mailman/listinfo/lustre-devel > -------------- next part -------------- An HTML attachment was scrubbed... URL: From gshipman at ornl.gov Tue Mar 8 21:55:13 2011 From: gshipman at ornl.gov (Shipman, Galen M.) Date: Tue, 08 Mar 2011 16:55:13 -0500 Subject: [Lustre-devel] LUG 2011 - One week left for early bird registration Message-ID: Early bird registration ends March 15th and our room block is quickly filling up. Register and book your room ASAP to ensure your spot! LUG 2011 will be held in Orlando, Florida from 8:30 AM Tuesday, April 12, 2011 through 12:00 noon April 14, 2011 at the Marriott World Center Golf and Spa resort. This two-and-a-half-day event is the primary venue for discussion and seminars on open source parallel file system technologies with a unique focus on the Lustre parallel file system. The conference is generously supported by the following corporate sponsors: Bull, DataDirect Networks, Dell, HP, LSI, Oracle, SGI, Terascala, Whamcloud, and Xyratex. The LUG program committee has put together a great agenda for this event, featuring presentations on Lustre features, upcoming enhancements, site-specific experiences, lessons learned, community organizations, and much more. You can find the entire LUG agenda via the LUG website at http://www.olcf.ornl.gov/event/lug-2011/ (click on the agenda tab) REGISTER TODAY You can register for LUG 2011 via the LUG website at http://www.olcf.ornl.gov/event/lug-2011/ (click on the registration tab) Early bird registration (through March 15) is $400 per person, while standard registration (after March 15) is $550 per person for the entire two-and-a-half-day event. Hotel reservations can be made using the website or phone number below: https://resweb.passkey.com/go/LUG2011 Tel: 1-800-266-9432 $179.00 (non-government) $104.00 (government or prevailing government rate) (Apologies if you are receiving this message more than once, it has been sent to multiple lists) From dmayfield at rgmadvisors.com Thu Mar 10 16:02:03 2011 From: dmayfield at rgmadvisors.com (Daniel Mayfield) Date: Thu, 10 Mar 2011 10:02:03 -0600 Subject: [Lustre-devel] Bug? Message-ID: I have a small test cluster (56 client nodes, 1 active MDS, 14 active OSSes). On Wednesday, I had a user run a job across a number of the 56 nodes, and we had a bunch of problems with filesystems/nodes hanging. On all of the server nodes I see a bunch of messages like this: -- odfs016.dmesg:Lustre: 6024:0:(lib-move.c:1826:lnet_parse_put()) Dropping PUT from 12345-192.168.50.238 at o2ib portal 28 match 1359131017746000 offset 0 length 368: 2 odfs016.dmesg:Lustre: 6020:0:(lib-move.c:1826:lnet_parse_put()) Dropping PUT from 12345-192.168.50.55 at o2ib portal 28 match 1359226700951811 offset 0 length 368: 2 odfs016.dmesg:Lustre: 6032:0:(lib-move.c:1826:lnet_parse_put()) Dropping PUT from 12345-192.168.50.56 at o2ib portal 28 match 1359136284378342 offset 0 length 368: 2 odfs016.dmesg:Lustre: 6033:0:(lib-move.c:1826:lnet_parse_put()) Dropping PUT from 12345-192.168.50.238 at o2ib portal 28 match 1359131017746019 offset 0 length 368: 2 odfs016.dmesg:Lustre: 6018:0:(lib-move.c:1826:lnet_parse_put()) Dropping PUT from 12345-192.168.50.55 at o2ib portal 28 match 1359226700951833 offset 0 length 368: 2 -- (for reference, .238 is the MDS+MGS, .55 and .56 are the last two client nodes). On some of the client nodes, I got actual OOPS messages: -- UG: unable to handle kernel NULL pointer dereference at (null) IP: [] sg_next+0x3/0x30 PGD 1839700067 PUD 18389e0067 PMD 0 Oops: 0000 [#1] SMP last sysfs file: /sys/devices/pci0000:00/0000:00:1e.0/0000:06:03.0/local_cpus CPU 1 Modules linked in: hidp l2cap bluetooth rfkill lmv mgc lustre lov osc lquota mdc fid fld ksocklnd ko2iblnd ptlrpc obdclass lnet lvfs libcfs rdma_ucm ib_sdp rdma_cm iw_cm ib_addr ib_ipoib ib_cm ib_sa ib_uverbs ib_umad iw_nes iw_cxgb3 cxgb3 ib_qib mlx4_ib mlx4_en mlx4_core ib_mthca ib_mad ib_core mptctl mptbase ipmi_devintf ipmi_si ipmi_msghandler dell_rbu netconsole configfs i2c_dev i2c_core nfs lockd fscache nfs_acl auth_rpcgss sunrpc ipv6 libcrc32c dca fuse ext3 jbd mbcache dm_mirror dm_multipath scsi_dh video output sbs sbshc acpi_pad parport_pc lp parport sg sr_mod cdrom bnx2 pata_acpi snd_pcm serio_raw ata_generic iTCO_wdt iTCO_vendor_support dcdbas snd_timer snd soundcore snd_page_alloc pcspkr dm_region_hash dm_log dm_mod ata_piix libata shpchp megaraid_sas sd_mod crc_t10dif scsi_mod xfs exportfs uhci_hcd ohci_hcd ssb mmc_core ehci_hcd [last unloaded: mlx4_core] Pid: 12935, comm: ptlrpcd-brw Not tainted 2.6.32.28-2.rgm #1 PowerEdge R610 RIP: 0010:[] [] sg_next+0x3/0x30 RSP: 0018:ffff88181c0c1760 EFLAGS: 00010246 RAX: 0000000000000000 RBX: ffff881823b19000 RCX: 0000000000000002 RDX: 0000000000000002 RSI: ffffc9001d3e7468 RDI: 0000000000000000 RBP: ffff88181c0c1800 R08: ffff8802d7555eb0 R09: 0000000000000002 R10: 0000000000000000 R11: ffff8802d7555eb0 R12: ffff880c38b26800 R13: ffff881823b14000 R14: ffff880c3946f090 R15: ffffc9001d3e7468 FS: 00007eff2e9f76e0(0000) GS:ffff880028200000(0000) knlGS:0000000000000000 CS: 0010 DS: 0000 ES: 0000 CR0: 000000008005003b CR2: 0000000000000000 CR3: 0000001836a84000 CR4: 00000000000006e0 DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000 DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7: 0000000000000400 Process ptlrpcd-brw (pid: 12935, threadinfo ffff88181c0c0000, task ffff88181c0be8c0) Stack: ffff88181c0c1800 ffffffffa094dc3e ffff880c0555c080 00050000c0a832fc <0> ffff88181c0c17d0 ffffffffa094e8a1 ffff88073f2bc140 ffff880c0555c118 <0> ffff880028333280 ffff880c00000002 ffffffff815cd3a0 0000000200000001 Call Trace: [] ? kiblnd_map_tx+0x1be/0x430 [ko2iblnd] [] ? kiblnd_queue_tx_locked+0x91/0x2b0 [ko2iblnd] [] kiblnd_setup_rd_iov+0x142/0x270 [ko2iblnd] [] kiblnd_send+0x5d9/0x970 [ko2iblnd] [] lnet_ni_send+0x51/0xd0 [lnet] [] lnet_send+0x5ab/0x960 [lnet] [] ? cfs_alloc+0x63/0x90 [libcfs] [] ? lnet_prep_send+0x50/0xb0 [lnet] [] LNetPut+0x2a7/0x7b0 [lnet] [] ? LNetMDBind+0x1c7/0x3f0 [lnet] [] ? cfs_alloc+0x63/0x90 [libcfs] [] ptl_send_buf+0x199/0x580 [ptlrpc] [] ? LNetMDAttach+0x350/0x4a0 [lnet] [] ptl_send_rpc+0x4be/0xc80 [ptlrpc] [] ptlrpc_send_new_req+0x3d8/0x810 [ptlrpc] [] ? find_busiest_queue+0x53/0x110 [] ? finish_task_switch+0x4f/0xa0 [] ptlrpc_check_set+0x308/0x19c0 [ptlrpc] [] ? try_to_del_timer_sync+0x7b/0xe0 [] ? del_timer_sync+0x22/0x30 [] ptlrpcd_check+0x1c0/0x250 [ptlrpc] [] ptlrpcd+0x363/0x3e0 [ptlrpc] [] ? default_wake_function+0x0/0x20 [] child_rip+0xa/0x20 [] ? cfs_alloc+0x63/0x90 [libcfs] [] ? ptlrpcd+0x0/0x3e0 [ptlrpc] [] ? child_rip+0x0/0x20 Code: 41 5d 41 5e 41 5f c9 c3 55 48 c7 c2 f0 c4 1e 81 be 80 00 00 00 48 89 e5 e8 6b ff ff ff c9 c3 66 0f 1f 84 00 00 00 00 00 55 31 c0 07 02 48 89 e5 75 0d 48 8b 57 20 48 8d 47 20 f6 c2 01 75 02 RIP [] sg_next+0x3/0x30 RSP CR2: 0000000000000000 ---[ end trace d0f44533e422125f ]--- BUG: unable to handle kernel NULL pointer dereference at (null) IP: [<(null)>] (null) PGD 182505b067 PUD 0 Oops: 0010 [#2] SMP last sysfs file: /sys/devices/pci0000:00/0000:00:1e.0/0000:06:03.0/local_cpus CPU 0 Modules linked in: hidp l2cap bluetooth rfkill lmv mgc lustre lov osc lquota mdc fid fld ksocklnd ko2iblnd ptlrpc obdclass lnet lvfs libcfs rdma_ucm ib_sdp rdma_cm iw_cm ib_addr ib_ipoib ib_cm ib_sa ib_uverbs ib_umad iw_nes iw_cxgb3 cxgb3 ib_qib mlx4_ib mlx4_en mlx4_core ib_mthca ib_mad ib_core mptctl mptbase ipmi_devintf ipmi_si ipmi_msghandler dell_rbu netconsole configfs i2c_dev i2c_core nfs lockd fscache nfs_acl auth_rpcgss sunrpc ipv6 libcrc32c dca fuse ext3 jbd mbcache dm_mirror dm_multipath scsi_dh video output sbs sbshc acpi_pad parport_pc lp parport sg sr_mod cdrom bnx2 pata_acpi snd_pcm serio_raw ata_generic iTCO_wdt iTCO_vendor_support dcdbas snd_timer snd soundcore snd_page_alloc pcspkr dm_region_hash dm_log dm_mod ata_piix libata shpchp megaraid_sas sd_mod crc_t10dif scsi_mod xfs exportfs uhci_hcd ohci_hcd ssb mmc_core ehci_hcd [last unloaded: mlx4_core] Pid: 22012, comm: java Tainted: G D 2.6.32.28-2.rgm #1 PowerEdge R610 RIP: 0010:[<0000000000000000>] [<(null)>] (null) RSP: 0018:ffff881094927460 EFLAGS: 00010097 RAX: ffff88181c0c1ee0 RBX: ffffffffffffffe8 RCX: 0000000000000000 RDX: 0000000000000000 RSI: 0000000000000003 RDI: ffff88181c0c1ee0 RBP: ffff8810949274a8 R08: 0000000000000000 R09: ffffffffa08c6e48 R10: ffff88182105c298 R11: 0000000000000001 R12: 0000000000000000 R13: ffff881837b31250 R14: 0000000000000000 R15: 0000000000000000 FS: 000000004030b940(0063) GS:ffff880c6a600000(0000) knlGS:0000000000000000 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 CR2: 0000000000000000 CR3: 00000014ef116000 CR4: 00000000000006f0 DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000 DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7: 0000000000000400 Process java (pid: 22012, threadinfo ffff881094926000, task ffff8811038d4680) Stack: ffffffff81040189 ffffffffa0a37b00 0000000300000001 ffff8810949275c8 <0> ffff881837b31248 0000000000000286 0000000000000003 0000000000000001 <0> 0000000000000000 ffff8810949274e8 ffffffff81047218 0000000000000000 Call Trace: [] ? __wake_up_common+0x59/0x90 [] __wake_up+0x48/0x70 [] cfs_waitq_signal+0x1a/0x20 [libcfs] [] ptlrpc_set_add_new_req+0x59/0x70 [ptlrpc] [] ptlrpcd_add_req+0x1e8/0x350 [ptlrpc] [] ? lustre_msg_get_opc+0x94/0x100 [ptlrpc] [] ? ptlrpc_lprocfs_brw+0xc0/0xd0 [ptlrpc] [] osc_send_oap_rpc+0x565/0xc10 [osc] [] ? cfs_mem_is_in_cache+0x16/0x60 [libcfs] [] ? cl_is_page+0x15/0x20 [obdclass] [] osc_check_rpcs+0x2a6/0x470 [osc] [] ? osc_page_transfer_add+0x4f/0x80 [osc] [] ? on_list+0x43/0x50 [osc] [] osc_io_submit+0x1b9/0x4c0 [osc] [] ? lov_attr_get+0x4a/0x80 [lov] [] cl_io_submit_rw+0x71/0x1a0 [obdclass] [] ? cfs_mem_is_in_cache+0x16/0x60 [libcfs] [] lov_io_submit+0x2be/0x960 [lov] [] cl_io_submit_rw+0x71/0x1a0 [obdclass] [] ? cl_page_top_trusted+0x16/0x60 [obdclass] [] cl_io_read_page+0xbe/0x1a0 [obdclass] [] ? cl_page_assume+0xa9/0x250 [obdclass] [] ll_readpage+0x98/0x1e0 [lustre] [] generic_file_aio_read+0x1e8/0x650 [] ? cl_object_attr_get+0x78/0x1a0 [obdclass] [] vvp_io_read_start+0x13b/0x3d0 [lustre] [] ? cl_wait+0xb5/0x260 [obdclass] [] cl_io_start+0x68/0x170 [obdclass] [] cl_io_loop+0x110/0x1d0 [obdclass] [] ? cl_env_info+0x15/0x20 [obdclass] [] ll_file_io_generic+0x242/0x3c0 [lustre] [] ? filemap_fault+0xd2/0x4b0 [] ? cl_env_get+0x29/0x330 [obdclass] [] ll_file_aio_read+0x1a0/0x2d0 [lustre] [] ? cl_env_get+0x180/0x330 [obdclass] [] ? cl_env_put+0x1b3/0x2e0 [obdclass] [] ll_file_read+0x171/0x310 [lustre] [] vfs_read+0xb5/0x1a0 [] sys_read+0x51/0x90 [] system_call_fastpath+0x16/0x1b Code: Bad RIP value. RIP [<(null)>] (null) RSP CR2: 0000000000000000 ---[ end trace d0f44533e4221260 ]--- -- Any help would be appreciated. daniel --------------------------------------------------------------- This email, along with any attachments, is confidential. If you believe you received this message in error, please contact the sender immediately and delete all copies of the message. Thank you. -------------- next part -------------- An HTML attachment was scrubbed... URL: From johann at whamcloud.com Thu Mar 10 17:13:24 2011 From: johann at whamcloud.com (Johann Lombardi) Date: Thu, 10 Mar 2011 18:13:24 +0100 Subject: [Lustre-devel] Bug? In-Reply-To: References: Message-ID: <20110310171323.GA2368@granier.hd.free.fr> On Thu, Mar 10, 2011 at 10:02:03AM -0600, Daniel Mayfield wrote: > Pid: 12935, comm: ptlrpcd-brw Not tainted 2.6.32.28-2.rgm #1 PowerEdge R610 ^^^^^^^^^^^ ^^^^^^^^^^^^^^^ Please note that 2.0 does not support 2.6.32 kernels. 2.1 will. > RIP: 0010:[] [] sg_next+0x3/0x30 > Call Trace: > [] ? kiblnd_map_tx+0x1be/0x430 [ko2iblnd] > [] ? kiblnd_queue_tx_locked+0x91/0x2b0 [ko2iblnd] This is a bug introduced by one of the patch from bug 21951 (attachment 29114). This patch was then reverted from both b1_8 (see 23123) and master (see 23332), but it is unfortunately included in 2.0.0. Cheers, Johann From gshipman at ornl.gov Tue Mar 15 03:27:15 2011 From: gshipman at ornl.gov (Shipman, Galen M.) Date: Mon, 14 Mar 2011 23:27:15 -0400 Subject: [Lustre-devel] LUG 2011 - One day left for early bird registration (save $150) Message-ID: <863E960F-76C1-4F3E-A49C-86DC50976DB8@ornl.gov> Early bird registration is only available through March 15th and our room block is nearly full. Register and book your room ASAP to ensure your spot! LUG 2011 will be held in Orlando, Florida from 8:30 AM Tuesday, April 12, 2011 through 12:00 noon April 14, 2011 at the Marriott World Center Golf and Spa resort. This two-and-a-half-day event is the primary venue for discussion and seminars on open source parallel file system technologies with a unique focus on the Lustre parallel file system. The conference is generously supported by the following corporate sponsors: Bull, DataDirect Networks, Dell, HP, LSI, Oracle, SGI, Terascala, Whamcloud, and Xyratex. The LUG program committee has put together a great agenda for this event, featuring presentations on Lustre features, upcoming enhancements, site-specific experiences, lessons learned, community organizations, and much more. You can find the entire LUG agenda via the LUG website at http://www.olcf.ornl.gov/event/lug-2011/ (click on the agenda tab) REGISTER TODAY You can register for LUG 2011 via the LUG website at http://www.olcf.ornl.gov/event/lug-2011/ (click on the registration tab) Early bird registration (through March 15) is $400 per person, while standard registration (after March 15) is $550 per person for the entire two-and-a-half-day event. Hotel reservations can be made using the website or phone number below: https://resweb.passkey.com/go/LUG2011 Tel: 1-800-266-9432 $179.00 (non-government) $104.00 (government or prevailing government rate) (Apologies if you are receiving this message more than once, it has been sent to multiple lists) From gshipman at ornl.gov Mon Mar 21 13:48:55 2011 From: gshipman at ornl.gov (Shipman, Galen M.) Date: Mon, 21 Mar 2011 09:48:55 -0400 Subject: [Lustre-devel] LUG 2011 - Last day to get a room in our room block (3/21/2011) Message-ID: The LUG 2011 room block expires today (3/21/2011). We have a limited number of rooms still available but we are almost full. If you plan on attending please reserve your room today! LUG 2011 will be held in Orlando, Florida from 8:30 AM Tuesday, April 12, 2011 through 12:00 noon April 14, 2011 at the Marriott World Center Golf and Spa resort. This two-and-a-half-day event is the primary venue for discussion and seminars on open source parallel file system technologies with a unique focus on the Lustre parallel file system. The conference is generously supported by the following corporate sponsors: Bull, DataDirect Networks, Dell, HP, LSI, Oracle, SGI, Terascala, Whamcloud, and Xyratex. The LUG program committee has put together a great agenda for this event, featuring presentations on Lustre features, upcoming enhancements, site-specific experiences, lessons learned, community organizations, and much more. You can find the entire LUG agenda via the LUG website at http://www.olcf.ornl.gov/event/lug-2011/ (click on the agenda tab) REGISTER TODAY You can register for LUG 2011 via the LUG website at http://www.olcf.ornl.gov/event/lug-2011/ (click on the registration tab) Standard registration is $550 per person for the entire two-and-a-half-day event. Hotel reservations can be made using the website or phone number below: https://resweb.passkey.com/go/LUG2011 Tel: 1-800-266-9432 $179.00 (non-government) $104.00 (government or prevailing government rate) (Apologies if you are receiving this message more than once, it has been sent to multiple lists) From vilobh.meshram at gmail.com Wed Mar 23 20:49:23 2011 From: vilobh.meshram at gmail.com (Vilobh Meshram) Date: Wed, 23 Mar 2011 16:49:23 -0400 Subject: [Lustre-devel] Problem while reading contents of file Message-ID: Hi, I have made some modifications in the Lustre source 1.8.3 . Now the thing is when I write to the file I can see in the OSS logs that the file size is getting updated properly but when I try to read some content of the file I cannot do that.... Also when I try to "list the content of the directory" with ls command I see ??? ??? ??? in the file attributes which are normally displayed by ls. Please let me know where I may be going wrong. Thanks, Vilobh -------------- next part -------------- An HTML attachment was scrubbed... URL: From nirant.puntambekar at oracle.com Thu Mar 24 18:00:11 2011 From: nirant.puntambekar at oracle.com (nirant.puntambekar at oracle.com) Date: Thu, 24 Mar 2011 11:00:11 -0700 (PDT) Subject: [Lustre-devel] Auto Reply: Lustre-devel Digest, Vol 60, Issue 10 Message-ID: <2dd71115-8cd2-43db-aad9-212e4b48ccb3@default> This is an auto-replied message. I am currently out of the office and will be back on March 30th. For anything urgent in my absence, do contact my manager Scott.Meadows at oracle.com. From aurelien.degremont at cea.fr Fri Mar 25 22:27:50 2011 From: aurelien.degremont at cea.fr (=?ISO-8859-1?Q?Aur=E9lien_Degr=E9mont?=) Date: Fri, 25 Mar 2011 23:27:50 +0100 Subject: [Lustre-devel] The good usage of lustre *_thread_info structure Message-ID: <4D8D16E6.8050708@cea.fr> Hello Doing some coding in Lustre, I'm wondering for a while was it the correct usage of thread_info structure like mdt_thread_info or mdd_thread_info. They contain pre-allocated data or pointer to pass this between function call and layer without overloading the stack. My concern is: if a function decide to use of them to store some of its data, how can it be sure that it was not used by an upper layer or a calling function? How can I be sure it is safe to use them? By example : struct mdt_thread_info { ... /* * Object attributes. */ struct md_attr mti_attr; ... } A function in MDT layer could decide it will use this structure (mti_attr) for its own need, then it will call several functions that could have the same need. How can those functions know that they can or cannot re-use this structure? Same issues for pointers. Thanks for any help Aurélien From di.wang at whamcloud.com Sat Mar 26 00:05:57 2011 From: di.wang at whamcloud.com (wangdi) Date: Fri, 25 Mar 2011 17:05:57 -0700 Subject: [Lustre-devel] The good usage of lustre *_thread_info structure In-Reply-To: <4D8D16E6.8050708@cea.fr> References: <4D8D16E6.8050708@cea.fr> Message-ID: <4D8D2DE5.30209@whamcloud.com> On 03/25/2011 03:27 PM, Aurélien Degrémont wrote: > Hello > > Doing some coding in Lustre, I'm wondering for a while was it the > correct usage of thread_info structure like mdt_thread_info or > mdd_thread_info. > They contain pre-allocated data or pointer to pass this between function > call and layer without overloading the stack. > My concern is: if a function decide to use of them to store some of its > data, how can it be sure that it was not used by an upper layer or a > calling function? > How can I be sure it is safe to use them? This thread_info will be initialized in the beginning of the request handling (ptlrpc_server_handle_request -> lu_context_init, and attached to the request), and currently the request will only be processed by a single thread, i.e. no other threads will try to access the request and the thread info. so it is safe to use this info within the service thread. This thread_info is actually designed for providing large temporary memory to functions, so they can get these memories in a cheap way, instead of allocating/freeing every time or reserving large tmp var in the stack. Each layer has its own thread info, it should not be accessed by different layers. Thanks WangDi > By example : > > struct mdt_thread_info { > ... > /* > * Object attributes. > */ > struct md_attr mti_attr; > ... > } > > A function in MDT layer could decide it will use this structure > (mti_attr) for its own need, then it will call several functions that > could have the same need. How can those functions know that they can or > cannot re-use this structure? Same issues for pointers. > > Thanks for any help > > Aurélien > > > > _______________________________________________ > Lustre-devel mailing list > Lustre-devel at lists.lustre.org > http://lists.lustre.org/mailman/listinfo/lustre-devel From aurelien.degremont at cea.fr Sat Mar 26 14:22:25 2011 From: aurelien.degremont at cea.fr (=?ISO-8859-1?Q?Aur=E9lien_Degr=E9mont?=) Date: Sat, 26 Mar 2011 15:22:25 +0100 Subject: [Lustre-devel] The good usage of lustre *_thread_info structure In-Reply-To: <4D8D2DE5.30209@whamcloud.com> References: <4D8D16E6.8050708@cea.fr> <4D8D2DE5.30209@whamcloud.com> Message-ID: <4D8DF6A1.8000004@cea.fr> Le 26/03/2011 01:05, wangdi said: > This thread_info will be initialized in the beginning of the request > handling (ptlrpc_server_handle_request -> lu_context_init, and attached > to the request), and currently the request will only be processed by a > single thread, i.e. no other threads will try to access the request and > the thread info. so it is safe to use this info within the service thread. > > This thread_info is actually designed for providing large temporary > memory to functions, so they can get these memories in a cheap way, > instead of allocating/freeing every time To be sure, what is the bad aspect we want to avoid here? Le 26/03/2011 08:52, Nikita Danilov said: > For the call chains within a layer it is up to the programmer to ascertain that the same element of lu_env::le_ctx is not re-used improperly. To some extent this can be done automatically, e.g., by introducing an interface to access sub-structures in struct mdt_thread_info, which would maintain a "busy" flag and check it on access. So it is rather a DIY :) If I take an example : mdd_thread_info.mti_xattr_buf is there to be used when reading/changing a XATTR. If I code an helper method in MDD which needs to manipulate an XATTR, this field looks really interesting. But, I do not know for sure that all callers we not be using it? So the only solution is that I verify each of them are not using it (and their callers)? And latter, if someone add a new way to call my method, he must looks for what stuff i'm using in the thread_info to be sure it does not conflict with what he currently uses? This seems a lot of code to check... And so, how can I chose that this field is relevant to be added to thread_info instead of allocating/freeing it inside my function? Thanks for both your help Aurélien From Nikita_Danilov at xyratex.com Sat Mar 26 07:52:49 2011 From: Nikita_Danilov at xyratex.com (Nikita Danilov) Date: Sat, 26 Mar 2011 00:52:49 -0700 Subject: [Lustre-devel] The good usage of lustre *_thread_info structure In-Reply-To: <4D8D16E6.8050708@cea.fr> References: <4D8D16E6.8050708@cea.fr> Message-ID: <73AED5C780AE05478241DB067651A92102000E2A@XYUS-EX22.xyus.xyratex.com> On Mar 26, 2011, at 01:27 , Aurélien Degrémont wrote: > Hello Hello Aurélien, > > Doing some coding in Lustre, I'm wondering for a while was it the > correct usage of thread_info structure like mdt_thread_info or > mdd_thread_info. > They contain pre-allocated data or pointer to pass this between function > call and layer without overloading the stack. > My concern is: if a function decide to use of them to store some of its > data, how can it be sure that it was not used by an upper layer or a > calling function? > How can I be sure it is safe to use them? the answer is: by trusting everybody to follow the convention. In the meta-data stack each "method" takes a pointer to struct lu_env is a parameter. This environment contains two "contexts" (struct lu_ctx): lu_env::le_ctx and lu_env::le_ses. By convention, lu_env::le_ctx should be used only as a substitute for an automatic variables, that is, it is used to conserve a space on the stack (which is a rather limited resource in kernel space) and should never me used to pass information between functions. lu_env::le_ses is a context embedded into a request, which can be used to store some request-global information, like credentials. For the call chains within a layer it is up to the programmer to ascertain that the same element of lu_env::le_ctx is not re-used improperly. To some extent this can be done automatically, e.g., by introducing an interface to access sub-structures in struct mdt_thread_info, which would maintain a "busy" flag and check it on access. Hope this helps, Nikita. > > By example : > > struct mdt_thread_info { > ... > /* > * Object attributes. > */ > struct md_attr mti_attr; > ... > } > > A function in MDT layer could decide it will use this structure > (mti_attr) for its own need, then it will call several functions that > could have the same need. How can those functions know that they can or > cannot re-use this structure? Same issues for pointers. > > Thanks for any help > > Aurélien > > > > _______________________________________________ > Lustre-devel mailing list > Lustre-devel at lists.lustre.org > http://lists.lustre.org/mailman/listinfo/lustre-devel ______________________________________________________________________ This email may contain privileged or confidential information, which should only be used for the purpose for which it was sent by Xyratex. No further rights or licenses are granted to use such information. If you are not the intended recipient of this message, please notify the sender by return and delete it. You may not use, copy, disclose or rely on the information contained in it. Internet email is susceptible to data corruption, interception and unauthorised amendment for which Xyratex does not accept liability. While we have taken reasonable precautions to ensure that this email is free of viruses, Xyratex does not accept liability for the presence of any computer viruses in this email, nor for any losses caused as a result of viruses. Xyratex Technology Limited (03134912), Registered in England & Wales, Registered Office, Langstone Road, Havant, Hampshire, PO9 1SA. The Xyratex group of companies also includes, Xyratex Ltd, registered in Bermuda, Xyratex International Inc, registered in California, Xyratex (Malaysia) Sdn Bhd registered in Malaysia, Xyratex Technology (Wuxi) Co Ltd registered in The People's Republic of China and Xyratex Japan Limited registered in Japan. ______________________________________________________________________ From Nikita_Danilov at xyratex.com Sat Mar 26 14:38:46 2011 From: Nikita_Danilov at xyratex.com (Nikita Danilov) Date: Sat, 26 Mar 2011 07:38:46 -0700 Subject: [Lustre-devel] The good usage of lustre *_thread_info structure In-Reply-To: <4D8DF6A1.8000004@cea.fr> References: <4D8D16E6.8050708@cea.fr> <4D8D2DE5.30209@whamcloud.com> <4D8DF6A1.8000004@cea.fr> Message-ID: <73AED5C780AE05478241DB067651A92102000E2C@XYUS-EX22.xyus.xyratex.com> On Mar 26, 2011, at 17:22 , Aurélien Degrémont wrote: > Le 26/03/2011 01:05, wangdi said: >> This thread_info will be initialized in the beginning of the request >> handling (ptlrpc_server_handle_request -> lu_context_init, and attached >> to the request), and currently the request will only be processed by a >> single thread, i.e. no other threads will try to access the request and >> the thread info. so it is safe to use this info within the service thread. >> >> This thread_info is actually designed for providing large temporary >> memory to functions, so they can get these memories in a cheap way, >> instead of allocating/freeing every time > To be sure, what is the bad aspect we want to avoid here? Kernel stack space is severely limited (4KB or 8KB) and dynamic memory allocation is too expensive for some paths. > > > Le 26/03/2011 08:52, Nikita Danilov said: >> For the call chains within a layer it is up to the programmer to ascertain that the same element of lu_env::le_ctx is not re-used improperly. To some extent this can be done automatically, e.g., by introducing an interface to access sub-structures in struct mdt_thread_info, which would maintain a "busy" flag and check it on access. > So it is rather a DIY :) > > If I take an example : > > mdd_thread_info.mti_xattr_buf > > is there to be used when reading/changing a XATTR. If I code an helper > method in MDD which needs to manipulate an XATTR, this field looks > really interesting. But, I do not know for sure that all callers we not > be using it? So the only solution is that I verify each of them are not > using it (and their callers)? And latter, if someone add a new way to > call my method, he must looks for what stuff i'm using in the > thread_info to be sure it does not conflict with what he currently uses? > This seems a lot of code to check... Well, c'est la vie. :-) I just ran a grep and found that ->mti_xattr_buf is used by * mdd_create(), * __mdd_lma_get(), * __mdd_lma_set(), * mdd_acl_chmod() and * mdd_check_acl(). Generally, foo_thread_info fields can only be used within the foo layer and one modifying it, is assumed to understand the layer. Thank you, Nikita. > > > And so, how can I chose that this field is relevant to be added to > thread_info instead of allocating/freeing it inside my function? > > > Thanks for both your help > > Aurélien > _______________________________________________ > Lustre-devel mailing list > Lustre-devel at lists.lustre.org > http://lists.lustre.org/mailman/listinfo/lustre-devel ______________________________________________________________________ This email may contain privileged or confidential information, which should only be used for the purpose for which it was sent by Xyratex. No further rights or licenses are granted to use such information. If you are not the intended recipient of this message, please notify the sender by return and delete it. You may not use, copy, disclose or rely on the information contained in it. Internet email is susceptible to data corruption, interception and unauthorised amendment for which Xyratex does not accept liability. While we have taken reasonable precautions to ensure that this email is free of viruses, Xyratex does not accept liability for the presence of any computer viruses in this email, nor for any losses caused as a result of viruses. Xyratex Technology Limited (03134912), Registered in England & Wales, Registered Office, Langstone Road, Havant, Hampshire, PO9 1SA. The Xyratex group of companies also includes, Xyratex Ltd, registered in Bermuda, Xyratex International Inc, registered in California, Xyratex (Malaysia) Sdn Bhd registered in Malaysia, Xyratex Technology (Wuxi) Co Ltd registered in The People's Republic of China and Xyratex Japan Limited registered in Japan. ______________________________________________________________________