[lustre-discuss] RHEL9: MDS LNet/ksocklnd issue after client reboot; lnetctl peer show -v hangs/crashes
Georgios Magklaras
georgiosm at met.no
Tue Jun 16 15:19:41 UTC 2026
We are seeing a serious LNet/ksocklnd issue on an active Lustre MDS/MGS and
would appreciate feedback from anyone who has seen similar behaviour.
Environment:
- Lustre version: 2.16.1
- OS: RHEL 9.4
- Kernel: 5.14.0-427.31.1_lustre.el9.x86_64
- Server role: active MDS/MGS
- LNet: TCP
Relevant module information:
filename:
/lib/modules/5.14.0-427.31.1_lustre.el9.x86_64/extra/lustre/fs/lustre.ko
version: 2.16.1
rhelversion: 9.4
vermagic: 5.14.0-427.31.1_lustre.el9.x86_64 SMP preempt mod_unload
modversions
Problem summary:
After some Lustre clients reboot, they are sometimes unable to remount the
filesystem. On the client side, the mount fails with:
mount.lustre: mount PRIMARYMDSIP at tcp:SECONDARYMDS at tcp:/APOLLO at
/lustre/metproductionB failed: Input/output error
Is the MGS running?
Also from the client:
lctl ping PRIMARYMDSIP at tcp
failed to ping PRIMARYMDSIP at tcp: Input/output error
On the active MDS/MGS, LNet still appears to have the expected local NI up:
net:
- net type: tcp
local NI(s):
- nid: PRIMARYMDSIP at tcp
status: up
interfaces:
0: enp65s0f0np0
statistics:
send_count: 3743678655
recv_count: 3802449073
drop_count: 1389
tunables:
peer_timeout: 180
peer_credits: 8
peer_buffer_credits: 0
credits: 256
lnd tunables:
conns_per_peer: 1
timeout: 49
However, when the system is in this bad state, running:
lnetctl peer show -v
on the active MDS does not just hang; in our experience it can crash the
system.
The only reliable recovery we have found so far is disruptive:
Unmount the MGS/MDT targets on the MDS.
Remove/reload the LNet/Lustre modules.
Remount the MGS/MDT targets.
After this, clients can mount again.
Representative MDS kernel messages:
LNet: Timeout error while writing to CLIENT1IP:1021. Closing socket: rc =
-110
We also see repeated ksocklnd / bulk I/O errors involving specific client
NIDs, for example from specific IPs that match 24.04.4 LTS (Noble Numbat)
2.16.1 clients that have the following form:
LNetError: socklnd_cb.c:1182:ksocknal_process_receive() [0000000011320d5e]
Error -71 on read from 12345-CLIENTIP1 at tcp ip CLIENTIP1:1021 LNetError:
socklnd.c:1580:ksocknal_destroy_conn() Completing partial receive from
12345-CLIENTIP1 at tcp[2], ip CLIENTIP1:1021, with error, wanted: 32768, left:
32768, last alive is 0 secs ago LNetError:
socklnd.c:1593:ksocknal_destroy_conn() Incomplete receive of lnet header
from 12345-CLIENTIP at tcp, ip CLIENTIP:1021, with error, protocol: 3.x.
LustreError: events.c:472:server_bulk_callback() event type 3, status -5
LustreError: ldlm_lib.c:3562:target_bulk_io() @@@ network error on bulk
WRITE LustreError: ldlm_lib.c:3556:target_bulk_io() @@@ Reconnect on bulk
WRITE
Basic IP reachability to the client NID may still work.
Clients that do not reboot/work.
Questions:
Has anyone seen similar behaviour with Lustre 2.16.1 on RHEL 9.4?
Are there known LU tickets involving lnetctl peer show -v hanging or
crashing in 2.16.x after client reconnect failures?
Are there known ksocklnd fixes after 2.16.1 that would make upgrading to
2.17.x advisable?
Is there a safer way to clear bad/stale LNet peer state on the MDS without
unloading/reloading LNet and remounting the MGS/MDT?
Are there specific diagnostics we should capture before recovery, apart
from SysRq blocked-task stacks and a vmcore?
Best regards,
GM
--
--
--
--
*Georgios Magklaras PhD*
Chief Engineer
IT Infrastructure/HPC
The Norwegian Meteorological Institute
https://www.met.no/
https://www.steelcyber.com/georgioshome/
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://lists.lustre.org/pipermail/lustre-discuss_lists.lustre.org/attachments/20260616/d1ce683b/attachment.html>
More information about the lustre-discuss
mailing list