Fio Workload Profiles for AI Agent Filesystem Benchmarking

Fio's three distinct modes expose hidden storage bottlenecks in AI workloads.

Features Editor · · 9 min read
Cover illustration for “Fio Workload Profiles for AI Agent Filesystem Benchmarking”
Storage Performance · October 2, 2026 · 9 min read · 2,117 words

A fio test that looks clean on paper can still leave a GPU cluster sitting idle. That is the uncomfortable truth behind most AI storage benchmarking today: fio builds independent I/O streams, one after another, each doing its own thing. Real training and agent workloads do not work that way. Storage operations inside a training run are locked in step with gradient computation and collective communication, which produces bursts of demand that a handful of independent fio streams simply does not reproduce.

That mismatch does not stay theoretical for long. Teams run fio, get a number, and provision storage against it, only to find GPU utilization in production has no relationship to the number they tested for. PRISM, published by the FAIR team at Meta AI, documents this concretely: sub-optimal storage usage on a notable cloud vendor produced a performance slowdown for a common data loading job, and a separate bug at a storage vendor caused a sharp increase in checkpoint latency. A standard fio run would have caught neither problem.

The reason runs deeper than a tuning mistake. AI agent workloads are not one I/O pattern wearing different clothes. They generate at least three distinct modes of access, each with its own block size, its own latency tolerance, and its own concurrency profile. Test them as a single blended workload and the result hides all three failure modes at once. The rest of this piece builds a profile for each mode, then shows how to run all three as one suite.

The three I/O modes an AI agent filesystem produces

An AI agent filesystem is really three workloads wearing a trench coat.

Mode 1 is the sequential read load that training data ingestion produces. Large files get pulled continuously at high bandwidth. The MLPerf Storage training workload specifies read requests in the ~170 KiB range, described as "very latency sensitive," with per-accelerator bandwidth targets that differ between A100 and H100 generations. Multiple processes run this load at once, one per GPU. The demand is not a single stream but many running in parallel, sustained across an entire training epoch rather than arriving in short bursts.

Mode 2 is the small random read pattern that retrieval-augmented generation and related agent work produce. Instead of reading through a corpus in order, retrieval jumps to specific chunks and embeddings scattered across storage, and it needs the answer back fast. GPU clusters doing inference or data ingestion depend on tight latency windows, and any delay in a metadata lookup compounds as it ripples through a training epoch. The block sizes here run orders of magnitude smaller than Mode 1, so a test tuned for sequential throughput reports numbers that look great while the real IOPS ceiling and latency spikes stay invisible.

Mode 3 is the large-block checkpoint write. Instead of a steady drip, this is a short, sharp burst where every rank in a distributed job writes its model and optimizer state at the same moment.

A filesystem can sail through Mode 1 and still buckle the instant Mode 3 hits, and the only way to find out ahead of time is to run both tests inside one suite rather than trusting either in isolation.

Building the sequential read profile for training data ingestion

Building the Mode 1 profile starts with block size, and the number is not arbitrary. The MLPerf Storage benchmark calls for reads around 170 KiB for its ResNet50-derived training workload, so a profile that quietly runs on fio's default block size is measuring something else: a workload different from the training job anyone actually cares about.

Concurrency comes next, set through numjobs. Multi-GPU nodes run one reader process per GPU, so numjobs needs to match the GPU count per node. CoreWeave's documented methodology for multi-server fio scaling treats this as the floor, not a nice-to-have, because anything less understates how many parallel readers are actually hammering the filesystem at once.

Direct I/O matters just as much. Setting --direct=1 forces reads to skip the OS page cache and hit the actual storage layer. Skip this flag and repeated reads get served out of DRAM, handing back a number that has nothing to do with how the storage system performs under real load.

For the I/O engine, libaio or io_uring both support async I/O, and io_uring (available on Linux 5.1 and later) cuts down syscall overhead once queue depth climbs. Queue depth itself should track how aggressively the real training framework prefetches data; a shallow iodepth setting understates demand compared to what a production data loader actually issues.

Runtime deserves attention too. Use --time_based with a runtime long enough for the filesystem to reach steady state, because a size-based test can stop before throughput settles on a warm filesystem, handing back a number that flatters the system under test.

A sample stanza pulls these together: rw=read, block size bracketing the MLPerf 170 KiB target, numjobs matching GPU count, ioengine=libaio, direct=1, iodepth=32, time_based, runtime=120. For file layout, SkyPilot's object-store benchmark methodology offers three useful templates to test against: one 500 GB file, eight 100 GB shards, or a 96-shard layout of 5 GB safetensors files.

What comes out the other end should center on sustained bandwidth in MiB/s as the headline number, with IOPS as a secondary check and p99 latency watched closely, since training loops stall the moment even a small fraction of reads run slow.

This profile tells a clear story about sustained throughput. It says nothing about what happens when every rank in the cluster writes a checkpoint at the same instant, which is exactly the gap Mode 3 fills.

Building the small random read profile for retrieval workloads

Retrieval workloads do not care about bandwidth. They care about how fast storage answers a question it was not expecting, and a profile built around throughput will award a passing grade to a filesystem that cannot actually hold up under real agent traffic.

Block size is the first place this profile diverges sharply from Mode 1. RAG chunk retrieval and metadata-heavy operations like git and tar work on small objects, so the test needs to run at small block sizes. Running it large instead inflates the bandwidth number because the IOPS ceiling and metadata latency problems stay hidden underneath.

The access pattern setting follows directly: rw=randread, because retrieval jumps to a specific chunk or key rather than reading in sequence. Concurrency should track the retrieval layer itself, not the GPU count, since numjobs here needs to reflect how many agent threads or processes are actually querying at once.

Latency reporting is where this profile earns its keep. p50, p95, p99, and p99.9 all need to show up in the output, and PRISM calls for latency-percentile reporting on this workload family for good reason: a result with a fine-looking average latency can still carry a p99 spike that signals metadata contention, and that contention interrupts an agent loop mid-task.

Queue depth is worth testing at both ends, since a low iodepth and a high iodepth model different retrieval behavior. A low iodepth of 1 to 4 models a synchronous retrieval call that blocks until it gets an answer, while a higher iodepth models async batch retrieval, and which one matters more depends on how the agent framework actually issues its requests.

Direct I/O still belongs in this profile (--direct=1), to keep the measurement on the filesystem rather than the kernel cache. Production retrieval working sets sometimes genuinely fit in cache, so running a second pass without direct=1 characterizes that cached case separately.

A real gap appears here: metadata-heavy operations, git clone and tar/untar among them, sit alongside data ingestion and checkpoint I/O as their own benchmark family in PRISM, but fio's randread mode cannot replicate them. Flag that gap rather than pretend fio covers it, and point toward PRISM and toward Google Cloud's published fio benchmark scripts for untar, git, and Python compilation workloads as a supplement.

The numbers that matter here are IOPS and p99 latency as the primary readouts, with bandwidth as a secondary note. A filesystem that cannot sustain the IOPS an agent needs, at a latency it can tolerate, will stall retrieval loops no matter how it performed on the sequential read test.

Building the checkpoint write profile for large-block burst writes

Checkpoint writes are where a filesystem's write ceiling gets exposed, because the write is large, every rank fires at once, and the whole cluster waits until it finishes.

Block size here runs large, matching the contiguous tensors a checkpoint actually writes. Testing with small blocks understates how efficiently the filesystem handles this kind of write, so the profile should mirror the real artifact shape rather than a generic small-write test. The access pattern is rw=write, sequential rather than random, since each shard of a checkpoint gets written straight through to its file.

Concurrency is the parameter that makes or breaks this profile. The checkpoint write rate is bounded by how many ranks are writing at the same moment, so numjobs has to match the real number of checkpointing processes in the workload under test. It can run into the dozens. That scale is why PRISM treats DDP/FSDP save and load as its own benchmark family rather than folding it into generic write tests.

Timing works differently here too. This is not a sustained-demand test like the training-read profile; it is a short, high-intensity burst, so the profile should run against a fixed size rather than a fixed runtime, with the test file sized to match a realistic checkpoint artifact, whether that means a single 500 GB file or one of SkyPilot's multi-shard layouts.

The flag most often left out, and the one that matters most, is fsync. Real checkpoints have to be durable on disk before training can resume, so add --fsync=1 or --fdatasync=1 to measure the full write-plus-flush latency rather than the time it takes to hand data to a buffer. That flush number is what actually determines how long the cluster sits idle waiting. Skipping it makes the profile report a bandwidth figure that looks great on a slide and has nothing to do with real stall time.

What to measure: aggregate write bandwidth across all jobs using group_reporting, total time to write the full checkpoint size, and fsync latency specifically. A filesystem that buffers writes and returns control quickly, but takes a long time to actually flush to durable storage, will show a bandwidth number that flatters it while real stall time stays hidden.

For a sense of scale, FarmGPU's MLPerf Storage v3.0 results, run on Solidigm storage under simulated B200 accelerator load, reached hundreds of GiB/s in aggregate throughput at sub-2ms per-read latency. That figure is a useful calibration point for judging where a given checkpoint write result lands relative to a system built for this kind of load.

The next step is running them together.

Assembling the three profiles into a single coherent fio job file

None of the three profiles answers the question that matters most on its own: what happens when all three modes hit the same filesystem close together, the way they do in a real training run. Running them as one sequenced job file is the only way to catch interaction effects that three separate test runs will never reveal.

Structurally, this means one [global] stanza up top carrying the settings shared across all three profiles: ioengine, direct=1, group_reporting, and output-format=json so results can get parsed and compared programmatically. Each job section underneath only needs to override what actually differs between modes, meaning bs, rw, numjobs, iodepth, and size or runtime. Separate the sections with new_group so fio reports metrics per section instead of blending everything into one aggregate number that hides which mode did what.

Point the whole file at a single target directory, using --directory or a filename that resolves to the mounted filesystem under test, so all three modes are exercising the exact same storage path rather than three different corners of the system. That detail sounds small. It is a test that reflects the real deployment, rather than one that quietly tests three unrelated storage configurations and calls it a benchmark.

Sequencing the runs matters as much as the parameters inside them. Start cold, with the checkpoint write profile first, before anything has been cached or warmed. Follow it with the sequential read profile, then the small random read profile, so the filesystem faces the same rough order of operations a real training and retrieval pipeline would put it through. A filesystem that aces each mode in isolation, run back to back against the same path in the order a production system would actually throw them at it, is the only kind of result worth trusting for capacity planning.

Sources

  1. PRISM: Evaluating POSIX Storage Systems for AI Research Workflows
  2. Benchmarking Storage for AI Workloads
  3. Storage Benchmarking: Distributed File Storage
  4. 1. fio - Flexible I/O tester rev. 3.42 — fio 3.42-115-gcd29 documentation

More in Storage Performance