Measuring Metadata Operation Throughput on Cloud Filesystems
Metadata operations dominate AI filesystem performance, not data throughput.

Adding storage servers increases bandwidth close to a straight line. Metadata doesn't work that way. The namespace that tracks where every file lives, who can open it, and what sits in which folder has to stay consistent across the whole system, and that dependency is what keeps metadata from scaling the way raw data transfer does. That asymmetry is the starting point for understanding AI filesystem performance: stat, readdir, open, create, and unlink calls aren't rare housekeeping tasks that happen between the real work. In an AI workload, they are most of the work.
Traces from Baidu AI Cloud back this up with hard numbers: metadata operations make up somewhere between two-thirds and nearly all filesystem requests, and they dominate how long I/O takes to finish. It's about how many times the system has to look something up, and how long each lookup takes, not about how much data moves through the pipe.
Two things about AI training push metadata this high. First, training datasets are made of huge piles of small files: images, text snippets, feature vectors. Each one needs its own lookup, so the metadata cost per byte transferred is high. Second, training at scale with 3D parallelism spreads work across many GPUs, and each checkpoint cycle writes hundreds or thousands of separate checkpoint files. Every one of those files needs its own namespace operation, so metadata contention stacks up with every checkpoint, not just once at the start.
There's a structural reason this can't be avoided through clever engineering alone. Research from SwitchFS (EuroSys '26) lays it out: before a single byte of data moves, a client has to resolve a path through the metadata service just to find out where the data lives. That lookup happens on every file access, no exceptions. When files are small, the lookup can cost more than the actual read.
How metadata bottlenecks translate into GPU idle time and slower research cycles
A slow metadata layer doesn't just make file operations sluggish. It stalls GPUs, stretches out checkpointing, and piles hours onto research timelines, which makes storage design a direct line item in both training cost and how fast a team can iterate.
Meta's engineering blog, published in July 2026, spells out the mechanism in simple terms. Training runs synchronize across GPUs in steps. If one GPU is stuck waiting on a storage response, every other GPU in that step waits too. A single slow metadata call doesn't just slow one machine. It drags the whole cluster to a stop until that one response comes back.
Meta's own legacy BLOB-storage system shows what this looks like in practice. The stack had piled up metadata layers over the years, a namelayer, a volumeslayer, a containerlayer, so a single object GET had to pass through several sequential metadata round trips, some of which crossed regions. Adding those delays together brought cumulative latency to hundreds of milliseconds. For a web page loading in the background, that kind of delay goes unnoticed. For a GPU cluster burning through compute budget while it waits, it's a direct hit to the bottom line.
Research velocity takes a separate hit, distinct from GPU idle time. As GPU clusters spread across more regions and datasets grow larger, Meta found researchers were regularly waiting hours just to get datasets ingested and moved into the region where training would actually run. That delay has nothing to do with GPU speed. It's a metadata and data-movement problem that slows experiments before a single training step starts.
Inference workloads carry their own version of this sensitivity, and it looks different from training. Inference needs high IOPS at latency measured in fractions of a millisecond. A filesystem built to move large checkpoint files quickly can still fail an inference workload outright, because checkpoint bandwidth and metadata rate are two different things to optimize for, and passing one test says nothing about the other.
Why existing cloud filesystem benchmarks miss most of what matters for AI workloads
Most teams evaluating cloud filesystems still lean on bandwidth numbers and aggregate throughput figures. That picture rarely includes metadata operation rates. A system can post great numbers on a spec sheet and still choke on the exact workload it's meant to run, because the test that would have caught the problem was never part of the evaluation.
Part of the issue is that object storage systems don't publish IOPS figures in the first place, and that's not an oversight. The concept doesn't map cleanly onto how object storage handles traffic: requests get distributed across queues rather than routed through fixed-capacity channels the way filesystem operations are. That means the numbers most vendors do publish can't be lined up side by side with filesystem metadata rates. Comparing them is like comparing a speedometer reading to a fuel gauge. Both tell you something, but not the same thing.
The practical cost of this gap becomes visible after the infrastructure decision is already made. Teams pick a storage system based on advertised bandwidth, deploy it, and only then discover that bandwidth was never the constraint they needed to test for. By the time that becomes clear, the cluster is already running, and the mismatch is a production problem instead of a planning decision.
Metadata throughput measures per-operation rates and what they reveal about system behavior
Metadata throughput is a family of per-operation rates, covering create, stat, open, readdir, unlink, and mkdir, and each one behaves differently under load. Reading them together, rather than adding them up, is what makes the measurement useful.
Each operation carries its own cost. Create and mkdir both have to contend for the same parent directory's metadata, which makes them sensitive to how many clients are hitting that directory at once. SwitchFS (EuroSys '26) names parent-directory contention under skewed access patterns as the single biggest scalability limit for metadata systems that update synchronously. Stat and readdir behave differently again: they're read-heavy operations, and their peak rate doesn't stack on top of the peak rate for create and mkdir. Open is its own case, since it combines permission checks with a location lookup, which makes it sensitive to how many layers a metadata lookup has to pass through, the same stacked-lookup cost Meta ran into with its legacy architecture.
That last point carries a warning for anyone reading benchmark results: a filesystem's total metadata throughput is not the sum of each operation's individual peak rate. Treating it that way produces projections that look great on a slide and fall apart the moment a real, mixed workload hits the system.
Peak rate alone also doesn't tell the full story. Because GPU training synchronizes at step boundaries, what matters most is how long the slowest responses take, specifically the p95 and p99 latency for metadata calls, not the average and not the median. A system can show a high average throughput number and still have a long tail of slow responses hiding underneath it. Those slow responses cause GPUs to sit idle, even while the headline average looks fine.
How to run a metadata throughput benchmark that produces comparable, useful results
Getting a usable number out of a metadata benchmark takes three things working together: the right tool, a workload shaped like the one the system will actually run, and a reporting process that captures latency spread and CPU cost rather than just a single peak figure. Skipping any one of those three makes the result misleading rather than useful.
MDTest is the standard tool for this kind of measurement. It's now bundled into the IOR suite, and it works by running concurrent streams of mkdir, stat, rmdir, creat, open, close, and unlink operations against a filesystem, built specifically to show the peak rate at which a system can handle each one. That makes it the natural starting point for anyone evaluating a filesystem for AI work, rather than a tool pulled in as an afterthought.
How MDTest gets configured changes what the result actually shows. Directory depth and file count matter a lot: a wide, shallow tree with millions of files stresses readdir and stat differently than a deep tree with only a handful of files per folder, and AI training datasets tend to produce exactly the wide, shallow kind. Concurrency level affects whether contention effects appear at all, since they only occur once enough clients are hitting the system at the same time; running the test single-threaded produces a number that says nothing about how the system will behave across a real multi-GPU job. Skew matters too. A uniform random access pattern conceals the hotspot problems that matter most, while a skewed pattern, many clients hammering the same directory, exposes the contention failures that actually break checkpoint-heavy training runs.
MDTest only covers the metadata side. FIO is the standard complement for the data path, and it's the tool most practitioners already trust for reporting IOPS, throughput, and latency. A complete benchmark reports both: MDTest's metadata numbers alongside FIO's data numbers, with p95 and p99 latency included rather than just an average.
CPU cost belongs in the report too. A system that hits a high metadata IOPS number by burning a disproportionate amount of CPU on the client or server side isn't delivering the same result as a system that hits that number efficiently. Reporting operation rates without reporting CPU utilization alongside them hides half the cost of the result.
Finally, the benchmark needs to cover the whole path a request actually takes, from the application layer, through the filesystem layer, across the network, down to the storage media itself. A number that stops at the filesystem API and ignores network round trips will look great in a test and fail to reproduce once the system is running in production.
Where current filesystem designs that address metadata throughput limits put the tradeoffs
Every approach that raises metadata throughput comes with a tradeoff attached, and knowing what each one gives up is what makes a benchmark result interpretable rather than just a number on a page.
The most direct fix is flattening the lookup path itself. Meta rebuilt its BLOB storage around a fat client SDK that pulls a read plan straight from the API server and streams bytes directly from storage servers, cutting out the layered proxy hops that made its old system too slow for AI training. The result: metadata became accessible within 1 to 2 milliseconds, with a high cache hit rate powered by distributed caching held in GPU host memory. Fewer hops means less latency, but it also means more logic has to live in the client itself.
A different fix decouples metadata capacity from storage capacity. Amazon FSx for Lustre added a scalable metadata feature that lets metadata IOPS scale up on its own, independent of how much storage is provisioned. In one documented setup, a 12 TiB file system configured with scaled metadata delivered substantially higher metadata IOPS than the same 12 TiB without the feature turned on. That's two resources that used to be locked together now able to move independently, which matters a lot for workloads that need more metadata headroom without paying for more storage they don't need.
SwitchFS, presented at EuroSys '26 (held April 27 to 30, 2026, in Edinburgh), takes a more structural swing at the problem. The research argues that synchronous metadata updates are overly cautious in the first place, since directories usually aren't read again right after they're written. By deferring directory updates and batching them together with help from a programmable switch that tracks directory state, SwitchFS hides latency and cuts down on contention. Under skewed workloads, the paper reports substantially higher throughput and markedly lower latency than Emulated-InfiniFS, and on real-world workloads, a large improvement in end-to-end throughput over CephFS. This is a research result, not a shipped product, but it points at asynchronous updates as a serious direction for the field.
Google Cloud Managed Lustre, announced in April 2026 at Google Next '26, removes the problem from software. Using TPUDirect and RDMA, data moves straight to accelerators without passing through the host CPU. Pulling the CPU out of the data path removes a source of jitter that no amount of software tuning could fix on its own. The announcement reported 10 TB/s of throughput, with performance well above other hyperscaler-managed Lustre offerings running as a single instance.
Pure object storage has no atomic rename operation, which forces applications to build their own workarounds. Systems that add rename support generally do it with inode-like structures, and those structures carry their own per-operation overhead. There's no way around this cost, only a choice about where to pay it, and any benchmark that runs a rename-heavy workload against an object-store-backed filesystem will show exactly where that cost lands.
What a metadata throughput result tells you about infrastructure fit for AI workloads
A metadata throughput number only means something once it's read against the specific operation mix, concurrency level, and latency spread of the workload it needs to serve. The same headline figure can describe a system that's a great fit for one job and a poor fit for another, depending on what that job actually asks the filesystem to do.
Training workloads with frequent checkpointing live or die on create and mkdir throughput under heavy concurrency, and on p99 latency specifically. A system that posts strong average stat rates but falls apart under directory contention will stall at the exact moments a training run can least afford it. Inference workloads ask for something different: high sustained IOPS and open and stat latency measured in fractions of a millisecond, since inference depends on fast random access into model weight files rather than on bulk throughput. Mixed research pipelines, the kind that combine dataset ingestion, training, and evaluation in one environment, benefit most from metadata capacity that can scale on its own, separate from storage volume, the same decoupling the FSx scalable metadata feature demonstrates. For that kind of pipeline, the ability to add metadata headroom without buying more storage affects reliability more directly than any single peak throughput number.
File aggregation deserves a place in this picture too, since it's a measurable way to shrink the problem rather than just benchmark around it. Packing many small files into fewer larger ones, through tar archives, WebDataset shards, or similar formats, cuts the number of metadata operations a training step has to issue. Running the same benchmark before and after aggregation shows how much of a system's metadata ceiling is really a dataset-layout problem dressed up as a filesystem limitation.
A specific architectural pattern applies when this benchmark data gets applied to real infrastructure choices: a filesystem that presents a standard POSIX interface over an existing object storage bucket, whether that's S3, GCS, R2, or Azure Blob, while caching on local NVMe with sub-millisecond latency on cache hits. That design routes repeated stat and open calls to the cache instead of sending them down to the object store's own metadata path, which is exactly where object storage tends to be weakest. For that kind of system, a cache-hit rate and a cache-miss latency figure tell a buyer far more than a single bandwidth number.


