Every filesystem benchmark you've ever read measures the wrong thing. They format a fresh partition, stream sequential writes, report MB/s, and call it a day. Meanwhile your production cluster is choking on snapshot metadata storms, silent corruption discovered during scrub, and fsync tail latency that makes databases stall. Bartosz Fenski's modern-fs-benchmark project finally tests what actually matters — and the results should terrify anyone running CoW filesystems at scale.

The benchmark suite runs 593 CI jobs on kernel 7.0.0-1012-azure, tracking 145 trend points across Btrfs, ZFS, and bcachefs with ext4 and XFS over md/LVM as baselines [1]. Crucially, it uses loop devices on shared ephemeral VMs — compare shapes and ratios, not absolute throughput. Each job records a host-calibration anchor. This isn't about who wins a drag race on empty hardware; it's about who survives the production gauntlet.

Classic benchmarks skip the workloads that cause 3 AM pages: redundancy layout changes under load, snapshot aging and scaling, transparent compression tradeoffs, encryption (native vs LUKS), reflink behavior, fsync tail latency, degraded operation and rebuild, corruption self-healing, and near-full ENOSPC behavior [1]. The project explicitly calls out that "if you keep data you care about on a non-checksumming stack, no benchmark number compensates for corruption you won't discover until years later." That risk is invisible in every classic benchmark; here it's a first-class result.

Phoronix's August 2024 run on an AMD EPYC 8534P with a Solidigm D7-PS1010 7.6 TB PCIe 5.0 NVMe tells the same story from a different angle. ZFS wasn't even included. Btrfs ranked last or second-to-last in all four workloads tested on Linux 6.11-rc2 [2]. The George Mason University billion-file study is even grimmer: ZFS took 92 hours to process one billion files, and Btrfs couldn't finish the read test at all [2].

XFS was the only filesystem requiring zero reconfiguration before the test could run. These aren't edge cases — they're structural limits.

The pain points map directly to your budget and on-call rotation. A 16 TB ZFS pool with deduplication enabled needs approximately 98 GB of RAM reserved for ZFS alone [2]. That's not a tuning parameter; it's a hardware tax. Btrfs is lighter and quicker to set up with no big cache to warm up, but heavy writes still hurt — especially on spinning disks and metadata-heavy workloads [3].

The fragmentation problem "is better than it used to be, but it still appears in database-style workloads where the same files are rewritten constantly" [3]. The nodatacow mount option helps, but it's a band-aid that sacrifices CoW guarantees.

Failure modes the vendor docs won't highlight: bcachefs is still maturing — its native encryption and erasure coding are compelling but unproven at scale. ZFS's ARC caching is brilliant until memory pressure forces eviction storms that cascade into latency spikes. Btrfs's RAID5/6 implementation remains "not recommended for production" after a decade. All three CoW filesystems amplify write amplification under snapshot-heavy workloads, and none of them handle ENOSPC gracefully when the allocator runs out of contiguous free space. The modern-fs-benchmark dashboard makes corruption + scrub a first-class metric precisely because silent data corruption is the failure mode that destroys businesses.

Your Monday morning blueprint: stop trusting vendor benchmarks and synthetic throughput numbers. Run modern-fs-benchmark's workload suite against your actual hardware topology — the repository is at github.com/fenio/modern-fs-benchmark [1]. Test your specific snapshot retention policy, your actual compression algorithm, your real encryption stack. Measure fsync p99 latency under load, not average throughput.

Verify corruption detection and self-healing with injected bit flips. Size RAM for ZFS deduplication before you enable it — 98 GB per 16 TB is not a typo [2]. If you're on Btrfs for database workloads, benchmark with nodatacow and without, then decide if the integrity tradeoff is worth the throughput. And for the love of uptime, run scrubs on a schedule — the benchmark proves corruption detection is a feature, not a given.

Sources

  1. modern-fs-benchmark: Continuous benchmarks for multi-device CoW filesystems
  2. File System Performance Comparison Statistics 2024
  3. Btrfs vs ZFS: Performance, Features & Use Cases