[lustre-discuss] The pre-allocation and parallel offset writing lead to massive ZFS space waste

David Schanzenbach davidls at hawaii.edu
Mon Jun 22 20:30:05 UTC 2026


Hi Homer,

Is it possible you are hitting this know bug with Lustre 2.15.7 + zfs?
https://jira.whamcloud.com/browse/LU-19193


David

On 6/22/2026 6:20 AM, Homer Li via lustre-discuss wrote:
> Hi All
>      Using rclone to read from object storage and write in parallel to
> multiple versions of Lustre 2.15 + ZFS causes some space waste. This
> issue does not occur with ZFS and any version of Lustre 2.12.
>
> In multi-threaded mode, multiple pwrite64 calls with different offsets
> are invoked, while single-threaded sequential write operations
> (detailed below) do not result in space amplification. The Lustre
> implementation causes issues with ZFS space allocation.
>
> Below are the commands to reproduce this issue:
>   rclone copy /tmp/test_file /mnt/lustre/zfs_index1/file1 --transfers 1
> --multi-thread-streams 32
> zfs dnode=648
>
>   rclone copy /tmp/test_file /mnt/lustre/zfs_index1/file2 --transfers 1
> --multi-thread-streams 1
> zfs dnode=839
>
> client$ ls -lhs /mnt/lustre/zfs_index1/file1 /mnt/lustre/zfs_index1/file2
> /mnt/lustre/zfs_index1/file1:
> total 1.5G
> 1.5G -rw-r--r-- 1 root root 1.0G Jun 22 20:57 test_file
>
> /mnt/lustre/zfs_index1/file2:
> total 994M
> 994M -rw-r--r-- 1 root root 1.0G Jun 22 20:57 test_file
>
> client$ xxh128sum /mnt/lustre/zfs_index1/file1/test_file
> /mnt/lustre/zfs_index1/file2/test_file
> 16024760318c6298efd1151033ad2e9f  /mnt/lustre/zfs_index1/file1/test_file
> 16024760318c6298efd1151033ad2e9f  /mnt/lustre/zfs_index1/file2/test_file
>
> oss$ umount -t lustre
> oss$ zfs mount -a
>
> oss$ ls -lihs ./0/d5/229 ./300000401/d3/227
> 648 1.5G -rw-rw-rw- 1 root root 1.0G Jun 22 20:57 ./0/d5/229
> 839 994M -rw-rw-rw- 1 root root 1.0G Jun 22 20:57 ./300000401/d3/227
>
>
> oss$ xxh128sum ./0/d5/229 ./300000401/d3/227
> 16024760318c6298efd1151033ad2e9f  ./0/d5/229
> 16024760318c6298efd1151033ad2e9f  ./300000401/d3/227
>
> What should have been mapped by a single L1 indirect block fractured
> into 256 L1 blocks and an additional L2 indirect block, triggering
> performance degradations and an ARC footprint explosion.
>
> A 4KB dblk completely fragments orderly sequential writes into a
> massive amount of tiny random I/Os. This not only leads to
> catastrophic space alignment and parity amplification on RAID-Z
> arrays, but also severely wastes ARC memory cache due to the metadata
> tree structure.
>
> oss$ zdb -ddddd test_ost1/test_ost1 648
>      Object  lvl   iblk   dblk  dsize  dnsize  lsize   %full  type
>         648    3   128K     4K  1.46G     512     1G  100.00  ZFS plain file
>                                                 192   bonus  System attributes
> dnode flags: USED_BYTES USERUSED_ACCOUNTED USEROBJUSED_ACCOUNTED SPILL_BLKPTR
> dnode maxblkid: 262143
> path /O/0/d5/229
> ...
> Spill block: 0:3b01ca7000:3000 200L/200P F=1 B=7452/7452
> cksum=000000046340a6b6:000001c2df0b4eca:00005d31f69d6ee9:000d32915fcfae51
> Indirect blocks:
>                 0 L2   0:3b01d04000:9000 20000L/3000P F=262144
> B=7452/7452 cksum=0000049e9f732fed:001a617da1cd2146:68ba901169b7e8d6:1203fa42b8e153ca
>                 0  L1  0:3b01caa000:6000 20000L/2000P F=1024
> B=7452/7452 cksum=0000019630118006:0008913b752091c4:18e62b360f606bea:31823a30bb201c6b
>                 0   L0 0:fc08b2000:3000 1000L/1000P F=1 B=7447/7447
> cksum=0000000000000000:0000000000000000:0000000000000000:0000000000000000
>              1000   L0 0:fc08b5000:3000 1000L/1000P F=1 B=7447/7447
> cksum=0000000000000000:0000000000000000:0000000000000000:0000000000000000
>              2000   L0 0:fc08b8000:3000 1000L/1000P F=1 B=7447/7447
> cksum=0000000000000000:0000000000000000:0000000000000000:0000000000000000
>              3000   L0 0:fc08bb000:3000 1000L/1000P F=1 B=7447/7447
> cksum=0000000000000000:0000000000000000:0000000000000000:0000000000000000
>              4000   L0 0:fc08be000:3000 1000L/1000P F=1 B=7447/7447
> cksum=0000000000000000:0000000000000000:0000000000000000:0000000000000000
>
>
>
> oss$ zdb -ddddd test_ost1/test_ost1 839
>      Object  lvl   iblk   dblk  dsize  dnsize  lsize   %full  type
>         839    2   128K     1M   994M     512     1G  100.00  ZFS plain file
>                                                 192   bonus  System attributes
> dnode flags: USED_BYTES USERUSED_ACCOUNTED USEROBJUSED_ACCOUNTED SPILL_BLKPTR
> dnode maxblkid: 1023
> path /O/300000401/d3/227
>
> ....
> Indirect blocks:
>                 0 L1  0:3b000e1000:6000 20000L/2000P F=1024 B=7452/7452
> cksum=00000229556c8638:000a9a2f615cde18:1d4cdd6e53eba13e:bbc5d925c17605f6
>                 0  L0 0:397f896000:201000 100000L/100000P F=1
> B=7449/7449 cksum=0000000000000000:0000000000000000:0000000000000000:0000000000000000
>            100000  L0 0:397fa97000:201000 100000L/100000P F=1
> B=7449/7449 cksum=0000000000000000:0000000000000000:0000000000000000:0000000000000000
>            200000  L0 0:397fc98000:201000 100000L/100000P F=1
> B=7449/7449 cksum=0000000000000000:0000000000000000:0000000000000000:0000000000000000
>            300000  L0 0:398009a000:201000 100000L/100000P F=1
> B=7449/7449 cksum=0000000000000000:0000000000000000:0000000000000000:0000000000000000
>            400000  L0 0:397fe99000:201000 100000L/100000P F=1
> B=7449/7449 cksum=0000000000000000:0000000000000000:0000000000000000:0000000000000000
>
>
>
> --
> May the wind guide your path; May the stars light your way
> _______________________________________________
> lustre-discuss mailing list
> lustre-discuss at lists.lustre.org
> https://urldefense.com/v3/__http://lists.lustre.org/listinfo.cgi/lustre-discuss-lustre.org__;!!PvDODwlR4mBZyAb0!TYRsnPPco3pZxVAJ_z0a9A3CCP3IIEmwfpDcJZwatwa929qir0JLvzxpHQQpFLmgbHriuLevycbJC-Q0DivC_EVuI8CF1E0$ 
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://lists.lustre.org/pipermail/lustre-discuss_lists.lustre.org/attachments/20260622/aefa2016/attachment.html>


More information about the lustre-discuss mailing list