[lustre-discuss] RHEL9: MDS LNet/ksocklnd issue after client reboot; lnetctl peer show -v hangs/crashes

Georgios Magklaras georgiosm at met.no
Tue Jun 16 15:19:41 UTC 2026


We are seeing a serious LNet/ksocklnd issue on an active Lustre MDS/MGS and
would appreciate feedback from anyone who has seen similar behaviour.

Environment:

   - Lustre version: 2.16.1
   - OS: RHEL 9.4
   - Kernel: 5.14.0-427.31.1_lustre.el9.x86_64
   - Server role: active MDS/MGS
   - LNet: TCP

Relevant module information:

filename:
/lib/modules/5.14.0-427.31.1_lustre.el9.x86_64/extra/lustre/fs/lustre.ko
version:        2.16.1
rhelversion:    9.4
vermagic:       5.14.0-427.31.1_lustre.el9.x86_64 SMP preempt mod_unload
modversions


Problem summary:

After some Lustre clients reboot, they are sometimes unable to remount the
filesystem. On the client side, the mount fails with:

mount.lustre: mount PRIMARYMDSIP at tcp:SECONDARYMDS at tcp:/APOLLO at
/lustre/metproductionB failed: Input/output error
Is the MGS running?

Also from the client:

lctl ping PRIMARYMDSIP at tcp
failed to ping PRIMARYMDSIP at tcp: Input/output error

On the active MDS/MGS, LNet still appears to have the expected local NI up:

net:
- net type: tcp
  local NI(s):
  - nid: PRIMARYMDSIP at tcp
    status: up
    interfaces:
      0: enp65s0f0np0
    statistics:
      send_count: 3743678655
      recv_count: 3802449073
      drop_count: 1389
    tunables:
      peer_timeout: 180
      peer_credits: 8
      peer_buffer_credits: 0
      credits: 256
    lnd tunables:
      conns_per_peer: 1
      timeout: 49

However, when the system is in this bad state, running:

lnetctl peer show -v

on the active MDS does not just hang; in our experience it can crash the
system.

The only reliable recovery we have found so far is disruptive:

Unmount the MGS/MDT targets on the MDS.
Remove/reload the LNet/Lustre modules.
Remount the MGS/MDT targets.

After this, clients can mount again.

Representative MDS kernel messages:

LNet: Timeout error while writing to CLIENT1IP:1021. Closing socket: rc =
-110

We also see repeated ksocklnd / bulk I/O errors involving specific client
NIDs, for example from specific IPs that match 24.04.4 LTS (Noble Numbat)
2.16.1 clients that have the following form:

LNetError: socklnd_cb.c:1182:ksocknal_process_receive() [0000000011320d5e]
Error -71 on read from 12345-CLIENTIP1 at tcp ip CLIENTIP1:1021 LNetError:
socklnd.c:1580:ksocknal_destroy_conn() Completing partial receive from
12345-CLIENTIP1 at tcp[2], ip CLIENTIP1:1021, with error, wanted: 32768, left:
32768, last alive is 0 secs ago LNetError:
socklnd.c:1593:ksocknal_destroy_conn() Incomplete receive of lnet header
from 12345-CLIENTIP at tcp, ip CLIENTIP:1021, with error, protocol: 3.x.
LustreError: events.c:472:server_bulk_callback() event type 3, status -5
LustreError: ldlm_lib.c:3562:target_bulk_io() @@@ network error on bulk
WRITE LustreError: ldlm_lib.c:3556:target_bulk_io() @@@ Reconnect on bulk
WRITE

Basic IP reachability to the client NID may still work.

Clients that do not reboot/work.

Questions:

Has anyone seen similar behaviour with Lustre 2.16.1 on RHEL 9.4?
Are there known LU tickets involving lnetctl peer show -v hanging or
crashing in 2.16.x after client reconnect failures?
Are there known ksocklnd fixes after 2.16.1 that would make upgrading to
2.17.x advisable?
Is there a safer way to clear bad/stale LNet peer state on the MDS without
unloading/reloading LNet and remounting the MGS/MDT?
Are there specific diagnostics we should capture before recovery, apart
from SysRq blocked-task stacks and a vmcore?

Best regards,

GM
-- 
-- 
--
--

*Georgios Magklaras PhD*
Chief Engineer
IT Infrastructure/HPC
The Norwegian Meteorological Institute

https://www.met.no/
https://www.steelcyber.com/georgioshome/
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://lists.lustre.org/pipermail/lustre-discuss_lists.lustre.org/attachments/20260616/d1ce683b/attachment.html>


More information about the lustre-discuss mailing list