<div dir="auto">Can you say more about these networking issues?<div dir="auto">Good to make a note of them in case anyone sees similar in the future. </div></div><br><div class="gmail_quote"><div dir="ltr" class="gmail_attr">On Fri, 12 May 2023, 20:40 Jane Liu via lustre-discuss, &lt;<a href="mailto:lustre-discuss@lists.lustre.org">lustre-discuss@lists.lustre.org</a>&gt; wrote:<br></div><blockquote class="gmail_quote" style="margin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex">Hi Jeff,<br>
<br>
Thanks for your response. We discovered later that the network issues <br>
originating from the iDRAC IP were causing the SAS driver to hang or <br>
experience timeouts when trying to access the drives. This resulted in <br>
the drives being kicked out.<br>
<br>
Once we resolved this issue, both the mkfs and mount operations started <br>
working fine.<br>
<br>
Thanks,<br>
Jane<br>
<br>
<br>
<br>
On 2023-05-10 12:43, Jeff Johnson wrote:<br>
&gt; Jane,<br>
&gt; <br>
&gt; You&#39;re having hardware errors, the codes in those mpt3sas errors<br>
&gt; define as &quot;PL_LOGINFO_SUB_CODE_OPEN_FAILURE_ORR_TIMEOUT&quot;, or in other<br>
&gt; words your SAS HBA cannot open a command dialogue with your disk. I&#39;d<br>
&gt; suspect backplane or cabling issues as an internal disk failure will<br>
&gt; be reported by the target disk with its own error code. In this case<br>
&gt; your HBA can&#39;t even talk to it properly.<br>
&gt; <br>
&gt; Is sdah the partner mpath device to sdef? Or is sdah a second failing<br>
&gt; disk interface?<br>
&gt; <br>
&gt; Looking at this, I don&#39;t think your hardware is deploy-ready.<br>
&gt; <br>
&gt; --Jeff<br>
&gt; <br>
&gt; On Wed, May 10, 2023 at 9:29\u202fAM Jane Liu via lustre-discuss<br>
&gt; &lt;<a href="mailto:lustre-discuss@lists.lustre.org" target="_blank" rel="noreferrer">lustre-discuss@lists.lustre.org</a>&gt; wrote:<br>
&gt; <br>
&gt;&gt; Hi,<br>
&gt;&gt; <br>
&gt;&gt; We recently attempted to add several new OSS servers ( RHEL 8.7 and<br>
&gt;&gt; Lustre 2.15.2). While creating new OSTs, I noticed that mdstat<br>
&gt;&gt; reported<br>
&gt;&gt; some disk failures after the mkfs, even though the disks were<br>
&gt;&gt; functional<br>
&gt;&gt; before the mkfs command. Our hardware admins managed to resolve the<br>
&gt;&gt; mdstat issue and restore the disks to normal operation. However,<br>
&gt;&gt; when I<br>
&gt;&gt; ran the mount OST command (when network had a problem and mount<br>
&gt;&gt; command<br>
&gt;&gt; timed out), similar problems occurred, and several disks were kicked<br>
&gt;&gt; <br>
&gt;&gt; out. The relevant /var/log/messages are provided below.<br>
&gt;&gt; <br>
&gt;&gt; This problem was consistent across all our OSS servers. Any insights<br>
&gt;&gt; <br>
&gt;&gt; into the possible cause would be appreciated.<br>
&gt;&gt; <br>
&gt;&gt; Jane<br>
&gt;&gt; <br>
&gt;&gt; -----------------------------<br>
&gt;&gt; <br>
&gt;&gt; May  9 13:33:15 sphnxoss47 kernel: LDISKFS-fs (md0): mounted<br>
&gt;&gt; filesystem<br>
&gt;&gt; with ordered data mode. Opts: errors=remount-ro<br>
&gt;&gt; May  9 13:33:15 sphnxoss47 systemd[1]: tmp-mntmirJ5z.mount:<br>
&gt;&gt; Succeeded.<br>
&gt;&gt; May  9 13:33:16 sphnxoss47 kernel: LNet: HW NUMA nodes: 2, HW CPU<br>
&gt;&gt; cores:<br>
&gt;&gt; 72, npartitions: 2<br>
&gt;&gt; May  9 13:33:16 sphnxoss47 kernel: alg: No test for adler32<br>
&gt;&gt; (adler32-zlib)<br>
&gt;&gt; May  9 13:33:16 sphnxoss47 kernel: Key type ._llcrypt registered<br>
&gt;&gt; May  9 13:33:16 sphnxoss47 kernel: Key type .llcrypt registered<br>
&gt;&gt; May  9 13:33:16 sphnxoss47 kernel: Lustre: Lustre: Build Version:<br>
&gt;&gt; 2.15.2<br>
&gt;&gt; May  9 13:33:16 sphnxoss47 kernel: LNet: Added LNI 169.254.1.2@tcp<br>
&gt;&gt; [8/256/0/180]<br>
&gt;&gt; May  9 13:33:16 sphnxoss47 kernel: LNet: Accept secure, port 988<br>
&gt;&gt; May  9 13:33:17 sphnxoss47 kernel: LDISKFS-fs (md0): mounted<br>
&gt;&gt; filesystem<br>
&gt;&gt; with ordered data mode. Opts:<br>
&gt;&gt; errors=remount-ro,no_mbcache,nodelalloc<br>
&gt;&gt; May  9 13:33:17 sphnxoss47 kernel: Lustre: sphnx01-OST0244-osd:<br>
&gt;&gt; enabled<br>
&gt;&gt; &#39;large_dir&#39; feature on device /dev/md0<br>
&gt;&gt; May  9 13:33:25 sphnxoss47 systemd-logind[8609]: New session 7 of<br>
&gt;&gt; user<br>
&gt;&gt; root.<br>
&gt;&gt; May  9 13:33:25 sphnxoss47 systemd[1]: Started Session 7 of user<br>
&gt;&gt; root.<br>
&gt;&gt; May  9 13:34:36 sphnxoss47 kernel: LustreError: 15f-b:<br>
&gt;&gt; sphnx01-OST0244:<br>
&gt;&gt; cannot register this server with the MGS: rc = -110. Is the MGS<br>
&gt;&gt; running?<br>
&gt;&gt; May  9 13:34:36 sphnxoss47 kernel: LustreError:<br>
&gt;&gt; 45314:0:(obd_mount_server.c:2027:server_fill_super()) Unable to<br>
&gt;&gt; start<br>
&gt;&gt; targets: -110<br>
&gt;&gt; May  9 13:34:36 sphnxoss47 kernel: LustreError:<br>
&gt;&gt; 45314:0:(obd_mount_server.c:1644:server_put_super()) no obd<br>
&gt;&gt; sphnx01-OST0244<br>
&gt;&gt; May  9 13:34:36 sphnxoss47 kernel: LustreError:<br>
&gt;&gt; 45314:0:(obd_mount_server.c:131:server_deregister_mount())<br>
&gt;&gt; sphnx01-OST0244 not registered<br>
&gt;&gt; May  9 13:34:39 sphnxoss47 kernel: Lustre: server umount<br>
&gt;&gt; sphnx01-OST0244<br>
&gt;&gt; complete<br>
&gt;&gt; May  9 13:34:39 sphnxoss47 kernel: LustreError:<br>
&gt;&gt; 45314:0:(super25.c:176:lustre_fill_super()) llite: Unable to mount<br>
&gt;&gt; &lt;unknown&gt;: rc = -110<br>
&gt;&gt; May  9 13:34:40 sphnxoss47 kernel: LDISKFS-fs (md1): mounted<br>
&gt;&gt; filesystem<br>
&gt;&gt; with ordered data mode. Opts: errors=remount-ro<br>
&gt;&gt; May  9 13:34:40 sphnxoss47 systemd[1]: tmp-mntXT85fz.mount:<br>
&gt;&gt; Succeeded.<br>
&gt;&gt; May  9 13:34:41 sphnxoss47 kernel: LDISKFS-fs (md1): mounted<br>
&gt;&gt; filesystem<br>
&gt;&gt; with ordered data mode. Opts:<br>
&gt;&gt; errors=remount-ro,no_mbcache,nodelalloc<br>
&gt;&gt; May  9 13:34:41 sphnxoss47 kernel: Lustre: sphnx01-OST0245-osd:<br>
&gt;&gt; enabled<br>
&gt;&gt; &#39;large_dir&#39; feature on device /dev/md1<br>
&gt;&gt; May  9 13:36:00 sphnxoss47 kernel: LustreError: 15f-b:<br>
&gt;&gt; sphnx01-OST0245:<br>
&gt;&gt; cannot register this server with the MGS: rc = -110. Is the MGS<br>
&gt;&gt; running?<br>
&gt;&gt; May  9 13:36:00 sphnxoss47 kernel: LustreError:<br>
&gt;&gt; 46127:0:(obd_mount_server.c:2027:server_fill_super()) Unable to<br>
&gt;&gt; start<br>
&gt;&gt; targets: -110<br>
&gt;&gt; May  9 13:36:00 sphnxoss47 kernel: LustreError:<br>
&gt;&gt; 46127:0:(obd_mount_server.c:1644:server_put_super()) no obd<br>
&gt;&gt; sphnx01-OST0245<br>
&gt;&gt; May  9 13:36:00 sphnxoss47 kernel: LustreError:<br>
&gt;&gt; 46127:0:(obd_mount_server.c:131:server_deregister_mount())<br>
&gt;&gt; sphnx01-OST0245 not registered<br>
&gt;&gt; May  9 13:36:08 sphnxoss47 kernel: Lustre: server umount<br>
&gt;&gt; sphnx01-OST0245<br>
&gt;&gt; complete<br>
&gt;&gt; May  9 13:36:08 sphnxoss47 kernel: LustreError:<br>
&gt;&gt; 46127:0:(super25.c:176:lustre_fill_super()) llite: Unable to mount<br>
&gt;&gt; &lt;unknown&gt;: rc = -110<br>
&gt;&gt; May  9 13:36:08 sphnxoss47 kernel: LDISKFS-fs (md2): mounted<br>
&gt;&gt; filesystem<br>
&gt;&gt; with ordered data mode. Opts: errors=remount-ro<br>
&gt;&gt; May  9 13:36:08 sphnxoss47 systemd[1]: tmp-mnt17IOaq.mount:<br>
&gt;&gt; Succeeded.<br>
&gt;&gt; May  9 13:36:09 sphnxoss47 kernel: LDISKFS-fs (md2): mounted<br>
&gt;&gt; filesystem<br>
&gt;&gt; with ordered data mode. Opts:<br>
&gt;&gt; errors=remount-ro,no_mbcache,nodelalloc<br>
&gt;&gt; Show less<br>
&gt;&gt; 11:03 AM<br>
&gt;&gt; <br>
&gt;&gt; -----------------------------<br>
&gt;&gt; <br>
&gt;&gt; it just repeats for all of the md raids, then the errors start and<br>
&gt;&gt; the<br>
&gt;&gt; drive fails and is disabled:<br>
&gt;&gt; <br>
&gt;&gt; May  9 13:44:31 sphnxoss47 kernel: LustreError:<br>
&gt;&gt; 48069:0:(super25.c:176:lustre_fill_super()) llite: Unable to mount<br>
&gt;&gt; &lt;unknown&gt;: rc = -110<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 kernel: mpt3sas_cm1:<br>
&gt;&gt; log_info(0x3112011a):<br>
&gt;&gt; originator(PL), code(0x12), sub_code(0x011a)<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 kernel: mpt3sas_cm1:<br>
&gt;&gt; log_info(0x3112011a):<br>
&gt;&gt; originator(PL), code(0x12), sub_code(0x011a)<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 kernel: mpt3sas_cm1:<br>
&gt;&gt; log_info(0x3112011a):<br>
&gt;&gt; originator(PL), code(0x12), sub_code(0x011a)<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 kernel: mpt3sas_cm1:<br>
&gt;&gt; log_info(0x3112011a):<br>
&gt;&gt; originator(PL), code(0x12), sub_code(0x011a)<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 kernel: mpt3sas_cm1:<br>
&gt;&gt; log_info(0x3112011a):<br>
&gt;&gt; originator(PL), code(0x12), sub_code(0x011a)<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 kernel: mpt3sas_cm1:<br>
&gt;&gt; log_info(0x3112011a):<br>
&gt;&gt; originator(PL), code(0x12), sub_code(0x011a)<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 kernel: mpt3sas_cm1:<br>
&gt;&gt; log_info(0x3112011a):<br>
&gt;&gt; originator(PL), code(0x12), sub_code(0x011a)<br>
&gt;&gt; ....<br>
&gt;&gt; ....<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 kernel: sd 16:0:31:0: [sdef] tag#1102<br>
&gt;&gt; FAILED<br>
&gt;&gt; Result: hostbyte=DID_SOFT_ERROR driverbyte=DRIVER_OK cmd_age=1s<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 kernel: sd 16:0:31:0: [sdef] tag#1102<br>
&gt;&gt; CDB:<br>
&gt;&gt; Read(10) 28 00 00 00 87 79 00 00 01 00<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 kernel: blk_update_request: I/O error,<br>
&gt;&gt; dev<br>
&gt;&gt; sdef, sector 277448 op 0x0:(READ) flags 0x84700 phys_seg 1 prio<br>
&gt;&gt; class 0<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 kernel: sd 16:0:31:0: [sdef] tag#6800<br>
&gt;&gt; FAILED<br>
&gt;&gt; Result: hostbyte=DID_SOFT_ERROR driverbyte=DRIVER_OK cmd_age=1s<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 kernel: sd 16:0:31:0: [sdef] tag#6800<br>
&gt;&gt; CDB:<br>
&gt;&gt; Read(10) 28 00 00 00 87 dd 00 00 01 00<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 kernel: blk_update_request: I/O error,<br>
&gt;&gt; dev<br>
&gt;&gt; sdef, sector 278248 op 0x0:(READ) flags 0x84700 phys_seg 1 prio<br>
&gt;&gt; class 0<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 kernel: device-mapper: multipath: 253:52:<br>
&gt;&gt; <br>
&gt;&gt; Failing path 128:112.<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 multipathd[6051]: sdef: mark as failed<br>
&gt;&gt; May  9 13:44:33 sphnxoss47 multipathd[6051]: mpathae: remaining<br>
&gt;&gt; active<br>
&gt;&gt; paths: 1<br>
&gt;&gt; ...<br>
&gt;&gt; ...<br>
&gt;&gt; May  9 13:44:34 sphnxoss47 kernel: mpt3sas_cm0:<br>
&gt;&gt; log_info(0x3112011a):<br>
&gt;&gt; originator(PL), code(0x12), sub_code(0x011a)<br>
&gt;&gt; May  9 13:44:34 sphnxoss47 kernel: mpt3sas_cm0:<br>
&gt;&gt; log_info(0x3112011a):<br>
&gt;&gt; originator(PL), code(0x12), sub_code(0x011a)<br>
&gt;&gt; May  9 13:44:34 sphnxoss47 kernel: mpt3sas_cm0:<br>
&gt;&gt; log_info(0x3112011a):<br>
&gt;&gt; originator(PL), code(0x12), sub_code(0x011a)<br>
&gt;&gt; May  9 13:44:34 sphnxoss47 kernel: md: super_written gets error=-5<br>
&gt;&gt; May  9 13:44:34 sphnxoss47 kernel: md/raid:md8: Disk failure on<br>
&gt;&gt; dm-55,<br>
&gt;&gt; disabling device.<br>
&gt;&gt; May  9 13:44:34 sphnxoss47 kernel: md: super_written gets error=-5<br>
&gt;&gt; May  9 13:44:34 sphnxoss47 kernel: md/raid:md8: Operation continuing<br>
&gt;&gt; on<br>
&gt;&gt; 9 devices.<br>
&gt;&gt; May  9 13:44:34 sphnxoss47 multipathd[6051]: sdah: mark as failed<br>
&gt;&gt; _______________________________________________<br>
&gt;&gt; lustre-discuss mailing list<br>
&gt;&gt; <a href="mailto:lustre-discuss@lists.lustre.org" target="_blank" rel="noreferrer">lustre-discuss@lists.lustre.org</a><br>
&gt;&gt; <a href="http://lists.lustre.org/listinfo.cgi/lustre-discuss-lustre.org" rel="noreferrer noreferrer" target="_blank">http://lists.lustre.org/listinfo.cgi/lustre-discuss-lustre.org</a> [1]<br>
&gt; <br>
&gt; --<br>
&gt; <br>
&gt; ------------------------------<br>
&gt; Jeff Johnson<br>
&gt; Co-Founder<br>
&gt; Aeon Computing<br>
&gt; <br>
&gt; <a href="mailto:jeff.johnson@aeoncomputing.com" target="_blank" rel="noreferrer">jeff.johnson@aeoncomputing.com</a><br>
&gt; <a href="http://www.aeoncomputing.com" rel="noreferrer noreferrer" target="_blank">www.aeoncomputing.com</a> [2]<br>
&gt; t: 858-412-3810 x1001   f: 858-412-3845<br>
&gt; m: 619-204-9061<br>
&gt; <br>
&gt; 4170 Morena Boulevard, Suite C - San Diego, CA 92117<br>
&gt; <br>
&gt; High-Performance Computing / Lustre Filesystems / Scale-out Storage<br>
&gt; <br>
&gt; Links:<br>
&gt; ------<br>
&gt; [1] <br>
&gt; <a href="https://urldefense.com/v3/__http://lists.lustre.org/listinfo.cgi/lustre-discuss-lustre.org__;!!P4SdNyxKAPE!B65twCaGe4aP1xnGrjpUnd-1OYuemL3X9zWyxfWEA54zk2tnvbhhrBFW5x9rXl7nFEkSsZpiRGIbodWHehLDQyvnK6u95iVHjg$" rel="noreferrer noreferrer" target="_blank">https://urldefense.com/v3/__http://lists.lustre.org/listinfo.cgi/lustre-discuss-lustre.org__;!!P4SdNyxKAPE!B65twCaGe4aP1xnGrjpUnd-1OYuemL3X9zWyxfWEA54zk2tnvbhhrBFW5x9rXl7nFEkSsZpiRGIbodWHehLDQyvnK6u95iVHjg$</a><br>
&gt; [2] <br>
&gt; <a href="https://urldefense.com/v3/__http://www.aeoncomputing.com__;!!P4SdNyxKAPE!B65twCaGe4aP1xnGrjpUnd-1OYuemL3X9zWyxfWEA54zk2tnvbhhrBFW5x9rXl7nFEkSsZpiRGIbodWHehLDQyvnK6vvMMT5RQ$" rel="noreferrer noreferrer" target="_blank">https://urldefense.com/v3/__http://www.aeoncomputing.com__;!!P4SdNyxKAPE!B65twCaGe4aP1xnGrjpUnd-1OYuemL3X9zWyxfWEA54zk2tnvbhhrBFW5x9rXl7nFEkSsZpiRGIbodWHehLDQyvnK6vvMMT5RQ$</a><br>
_______________________________________________<br>
lustre-discuss mailing list<br>
<a href="mailto:lustre-discuss@lists.lustre.org" target="_blank" rel="noreferrer">lustre-discuss@lists.lustre.org</a><br>
<a href="http://lists.lustre.org/listinfo.cgi/lustre-discuss-lustre.org" rel="noreferrer noreferrer" target="_blank">http://lists.lustre.org/listinfo.cgi/lustre-discuss-lustre.org</a><br>
</blockquote></div>