Benchmarking POSIX Filesystem Layers Over Object Storage

Standard benchmarks miss the bursty, mixed I/O patterns that actually slow AI training at scale.

Staff Writer · · 10 min read
Cover illustration for “Benchmarking POSIX Filesystem Layers Over Object Storage”
Storage Performance · October 6, 2026 · 10 min read · 2,175 words

The hard part isn't measuring these systems. The tools teams grab off the shelf were built to answer a different question, so they quietly skip past the failures that actually show up in production, and those failures don't arrive in isolation. That's a fine way to test a parallel file system when it's moving one big sequential stream. It's a bad way to test what happens when a researcher kicks off a training job, opens three Jupyter notebooks, and runs a data validation script at the same time, all hitting the same storage layer in bursts nobody scheduled.

Researchers at FAIR, Meta, spelled this gap out in a paper called PRISM, published in July 2026. Their argument: existing benchmarks judge storage systems only on peak performance, and in doing so, they miss the bursty, mixed-up I/O patterns that actually define how AI research gets done day to day. Real research workflows don't look like a sustained stream of identical requests. They look like a half-finished data pipeline, a checkpoint save interrupting a read job, a thousand small files opened for a sanity check, then abandoned an hour later. If a benchmark is tuned for sequential throughput, it will crown a storage system a winner right before that system falls over the moment someone opens an interactive session next to a training run.

The cost of that mismatch isn't abstract. Infrastructure teams buy based on peak-throughput leaderboards, roll the system into production, and only then run into the real problems: metadata storms, random writes that don't behave the way anyone modeled, checkpoint latency that spikes for no reason the dashboard can explain. By the time anyone notices, a cluster full of expensive GPUs is sitting at a fraction of the utilization it was bought for, waiting on storage instead of crunching numbers. PRISM's own findings put a number on this. Running a storage system sub-optimally on a major cloud vendor produced a substantial slowdown on an ordinary data loading job. Separately, a bug in a storage vendor's system drove a steep increase in checkpoint latency. Neither of those appears on a peak-throughput chart, because peak-throughput charts aren't built to look for them.

What the POSIX-to-object translation requires

POSIX and object storage are built on incompatible assumptions rather than being two flavors of the same thing with different speed dials, and any layer that translates between them has to decide, quietly, which guarantees it's willing to give up. A benchmark that only measures outcomes, without checking which semantic corners got cut to produce them, will miss exactly the gaps that matter.

Atomic rename, symlinks, hard links, in-place random writes, directory listing: none of these are rare edge cases that only matter to unusual programs. They sit inside ordinary application code, so nobody thinks about them until they break. They break because object storage has no cheap, native way to do what POSIX expects as a matter of course.

Metadata is the part of this that's easiest to miss and most expensive to ignore. A stat call, a rename, a chmod: none of these have a native equivalent in S3. Each one becomes a network round trip in place of a local lookup. Running a directory listing at training scale (the equivalent of typing ls on a folder with a few million files in it) fires off thousands of LIST requests per epoch, each one adding latency that a local file system would never charge.

Full POSIX compliance is a specific list, not a vague standard: atomic rename, flock/fcntl locking, mmap, fsync, hard links, symlinks, and sparse file support. Whether a given translation layer implements all of these, some of them, or an approximation of them determines something very concrete: whether existing code runs as written, or whether someone has to go back and rewrite the parts that assumed POSIX guarantees that were never actually there. It's tempting to assume most workloads never touch these sharper edges. They do, just not in ways anyone labels up front. Checkpoint recovery depends on atomic rename. Data preprocessing pipelines depend on symlinks and directory semantics. Ordinary developer tools, git, conda, tar, all depend on these same primitives without ever announcing the dependency.

How the three failure modes cluster under real workloads

Diagram: The Three Failure Modes That Compound Under Real Workloads. Visualizes: Show how three distinct storage failure modes — metadata contention, random write amplification, and the cost of faking atomic rename — interact and feed each other…

Real benchmarks against POSIX-over-object layers turn up the same three trouble spots every time: metadata contention, random write amplification, and the cost of faking atomic rename. They don't show up in isolation. They compound.

Metadata contention hits hardest in computer vision and multimodal workloads, the ones built on millions of small files: JPEGs, audio clips, short text snippets. In these cases the file system spends more time finding and opening files than it spends actually reading their contents. Bandwidth almost never limits these workloads. It doesn't matter how fast the pipes are if the bottleneck is the lookup, not the transfer. PRISM frames this directly: AI research workflows are metadata-heavy and latency-sensitive, and a benchmark built around peak throughput ends up actively misleading anyone using it to choose infrastructure.

Random write amplification comes from a different source: object stores treat objects as immutable. Changing four kilobytes in the middle of a file means the system can't just patch those four kilobytes. It has to read the whole object, modify it, and write the whole thing back, turning a small in-place edit into a full-object PUT that might span many megabytes. FUSE-based layers built on this model slow to a crawl on any workload that assumes files can be edited in place, and model checkpointing is exactly that kind of workload. Checkpoints write large tensors that partially overlap from one save to the next, and teams routinely underestimate how much this costs. Model and optimizer state can run to hundreds of gigabytes per save. When that checkpoint traffic shares an I/O path with ongoing training reads, training throughput drops.

Atomic rename rounds out the trio. On object storage, a rename is really a copy followed by a delete, two separate operations with a gap between them. That gap is where crash safety goes to die. Any code that relies on rename as an atomic swap, including PyTorch's standard checkpoint pattern of writing to a temp file and calling os.replace(tmp, final), inherits a consistency risk that didn't exist on a real POSIX file system. This is also why modern table formats like Apache Iceberg, Delta Lake, and Apache Hudi exist: to provide ACID transactions and consistent metadata management on top of object storage, problems that go well beyond the missing rename but include it. If a POSIX layer emulates rename as copy-then-delete, it carries the same exposure that those table formats were built to solve.

A benchmarking study out of NCAR, built around the Pangeo software stack, put numbers on how much this interacts with file format choice. Comparing Zarr on object storage against NetCDF on POSIX file systems, the study found throughput swung heavily depending on the pairing: Zarr delivered substantially better write throughput on POSIX, while NetCDF on POSIX beat Zarr on S3 by a wide margin on reads. The three failure modes don't sit in separate boxes, either. A workload that strains metadata handling is often the same workload moving files between pipeline stages, which triggers renames. Checkpoint writes can balloon from read-modify-write cycles, and when they land on the same storage path as ongoing reads, one failure mode feeds the next.

What PRISM's methodology reveals about benchmarking AI research storage

PRISM's real contribution is the method: testing a storage system across the entire research lifecycle, not just the high-load training phase, separates a benchmark that tells the truth from one that flatters a vendor's peak numbers. Published by Adithya Kumar, Aditya Basu, Jacob Kahn, Parth Malani, Leo Huang, and Kalyan Saladi at FAIR, Meta, the framework reproduces representative AI research workloads across data ingestion, checkpoint I/O, and developer workflows, scoring POSIX storage systems on usability and performance together, on real GPU clusters.

The lifecycle PRISM covers is wider than most benchmarks bother with: cloning git repositories, building packages and conda environments, pulling data from a remote dataset hub with curl, generating synthetic data, loading data for pretraining, validating datasets, preprocessing, and managing checkpoints. Not just the training loop. The whole mess around it.

One finding stands out because it cuts against conventional wisdom: a flash-backed NFS setup beat a flash-backed Lustre setup by a wide margin on the distributed checkpoint load use case. Parallel file systems like Lustre are the default assumption for GPU cluster storage, so a result like this says the default assumption needs checking, not blind trust. A benchmark that only measured write throughput would have picked Lustre. A benchmark that only measured checkpoint recovery would have picked NFS. A benchmark covering both operations at once reveals that the tradeoff exists.

That's the practical lesson for anyone evaluating a POSIX-over-object layer: the benchmark has to include the developer workflow and the interactive phase alongside the training phase. Metadata contention and rename failures occur precisely in those phases, the ones most benchmarks treat as an afterthought.

What MLPerf Storage v3.0 measures

MLPerf Storage v3.0 widens the standard benchmark's scope in a real way, but adding S3 object storage access next to POSIX doesn't yet close the gap between measuring peak performance and measuring what research workloads actually look like. MLCommons announced the v3.0 results on September 1, 2026. The new version aims to represent the full breadth of storage workloads AI systems generate, and for the first time it includes S3 object storage access alongside the existing POSIX measurement path.

Nineteen organizations submitted results this round, eleven of them first-time entrants, including Azure, NVIDIA, and Nebius. That's a meaningfully larger field than prior rounds produced, so buyers have more to compare against. Of those submissions, roughly one in six used the S3 storage access layer, a real foothold for object storage in the benchmark but not yet the dominant mode.

Everpure's FlashBlade//EXA took first place across the large parameter model checkpointing and KV cache categories, posting the leading write bandwidth for a very large parameter model across 30 data nodes. ZettaLane's submission stood out for a different reason: it was the only entry running Lustre directly on cloud object storage, delivering 32.42 GB/s of checkpoint write throughput for a Llama 3 model from two clients. That's a concrete sign that the POSIX-over-object category now has a seat at the table in the industry's standard benchmark, not just in academic papers.

What v3.0 still doesn't capture is the same gap PRISM was built to close: the bursty, developer-driven, interactive I/O that defines actual research use. MLPerf Storage stays oriented around peak throughput inside well-defined training and checkpointing phases. That's a meaningful measurement, but it's not the whole picture. A system that tops the MLPerf Storage charts can still buckle under the metadata-heavy, latency-sensitive patterns that occur once real researchers, not synthetic benchmark loads, start using it.

How the object storage tier shapes POSIX layer performance

A POSIX layer's performance depends partly on what the object storage underneath it actually guarantees, before the translation code does anything. If you run the same POSIX layer against two different object storage backends, the comparison stops being fair, because the floor each one stands on is different.

Latency is the clearest example. Standard S3 requests run in the hundreds of milliseconds for the first byte. S3 Express One Zone cuts that down to single-digit milliseconds. That gap is large enough to change which failure mode occurs first in any POSIX layer sitting on top, whether it's metadata lookups stacking up or checkpoint writes stalling.

Consistency matters just as much. S3 moved to strong consistency in December 2020, which was a precondition for any POSIX layer that needs read-after-write correctness for checkpoint recovery to actually work. Any benchmark run before that date measured a world that no longer exists for current deployments.

Directory buckets, introduced in 2023, cut the cost of LIST operations that made early FUSE layers so expensive for metadata-heavy work. If you benchmark the same POSIX translation code against a flat-namespace bucket, then against a directory bucket, the metadata numbers come out different, even though not a single line of the POSIX layer's code changed. The object storage tier set the floor; the translation layer just inherited it.

CoreWeave's AI Object Storage, expanded on October 16, 2025 after its original launch on March 20, 2025, makes the same point from a different angle. It routes throughput through a multi-cloud networking backbone built on private interconnects, direct cloud peering, and high-capacity ports, paired with a caching layer called the Local Object Transport Accelerator (LOTA), with InfiniBand networking as one of the technologies involved in keeping data close to compute. That caching sits in the object storage tier itself, below any POSIX translation layer running on top of it. Whatever POSIX layer gets bolted on will inherit that performance profile. The object storage tier is an active variable beneath any POSIX layer, and no benchmark result means much until it accounts for which object storage tier was underneath it.

Sources

  1. PRISM: Evaluating POSIX Storage Systems for AI Research Workflows
  2. Pangeo Benchmarking Analysis: Object Storage vs. POSIX File System Haiying Xu

More in Storage Performance