Multi Cloud Object Storage Performance Comparison for AI Data Pipelines
Storage bottlenecks—not algorithms—determine how fast your GPUs actually train.

AI pipelines stall for a boring reason: storage can't feed GPUs fast enough. That's the whole problem, and it's architectural, not algorithmic. No amount of clever modeling fixes it, because the storage layer feeding the model is the bottleneck.
Here's the mechanism. A data loader hands shards of training data to the optimizer. The GPU chews through a step, finishes, and waits for the next shard to show up. One wait might cost milliseconds. Multiply that across millions of training steps over a multi-week run; the wasted time creates a real budget problem and a real schedule problem. A single AI accelerator costs tens of thousands of dollars. An idle GPU isn't a technical inconvenience, it's a line item.
Training workloads also demand latency that most storage systems were never built to hit: sub-millisecond response times on small-object reads. That requirement pushed the industry toward all-flash parallel file systems, because that's what it takes to run fast fabric-based networking and keep up with it.
A second mismatch compounds that one. AI pipelines don't fit neatly into the file world or the object world. File-based systems produce the data upstream, but it then needs to feed downstream workloads that expect high-throughput object reads. Training and inference pull in different directions, too. Training wants sustained throughput for reading datasets and bursts of write capacity for checkpoints. Inference wants low latency for loading models, pulling KV-cache entries, and retrieving embeddings. If a generic object store sits in front of that inference workload without a fast metadata tier, embedding round-trips blow past any latency budget a real-time application can tolerate.
None of this means the answer is simply buying faster storage. The right fix depends on which stage of the pipeline is actually struggling, and that's the question the rest of this piece works through.
The three metrics that gate GPU utilization
Storage vendors love a big throughput number on a spec sheet. Most of those numbers can't tell you how the system will behave when real training jobs hit it. Sustained throughput, metadata latency on small objects, and co-location fit between storage and compute are what keep GPUs busy or leave them idle.
Sustained throughput matters more than peak throughput. A system that posts an impressive burst number but degrades once multiple nodes hit it at the same time will stall a distributed training job, no matter how good its spec sheet reads. Google Cloud found that Rapid Bucket cuts blocked GPU time by 50% when it runs multi-modal training. That number matters because it holds across a full epoch of sustained throughput, not just a short burst under ideal conditions.
Metadata latency on small objects is the second gate, and it is easy to miss because bulk-transfer benchmarks don't measure it. Training pipelines lean hard on listing, stat, and open calls across huge numbers of small files: tokenized shards, embedding indices, and similar. Those calls dominate the time a job spends starting up and the time it spends recovering from a checkpoint. A system can be excellent at sequential bulk reads and still bottleneck every single run at the start and the end, simply because metadata operations are slow.
Co-location fit is the third. Physical or logical distance between storage and compute reintroduces an egress path, and that path throttles throughput and adds latency that's hard to predict. A storage system can carry an impressive rated bandwidth number, but if it can't sit close to the GPU cluster it's feeding, it will still behave slowly in practice. Rated bandwidth describes a lab condition. Distance describes reality.
Taken together, these three metrics are less a shopping list than a filter. If you run any storage option through sustained throughput, metadata latency, and co-location fit, you can watch most marketing claims hold up or fall apart fast.
How egress fees and multi-cloud sprawl tax pipeline performance
Throughput and latency aren't the only things that gate performance. Cost does too, and it does so in a sneaky way: teams make architecture decisions to dodge a bill, not because the decision makes the pipeline faster. That tax on performance appears in architecture decisions, not in a benchmark.
Consider what it costs to move a large training corpus. If you move a large training dataset, or a large multimodal corpus, at standard cloud rates, a single move can run tens of thousands of dollars in egress fees. Training is iterative. Teams don't move data once, they move it repeatedly as pipelines get rebuilt and experiments get rerun, and each move multiplies that figure.
Multi-cloud setups make the tax worse. Most enterprises now run infrastructure across more than one cloud, and each cloud you add bolts on its own IAM model, its own billing system, and its own observability stack. Some practitioners report infrastructure bills climbing substantially from transfer fees and the overhead of managing all those separate systems, apart from any compute cost.
Pricing structures reward exactly the behavior that performance punishes. Both Azure and GCP charge noticeably more to move data than to store it. The cheapest path for a cloud bill is often the slowest path for a training job. Teams respond to that pricing logic by duplicating datasets across regions or clouds instead of paying egress fees to centralize them. Teams avoid egress fees, but datasets then drift out of sync across copies, synchronization introduces its own delay, and storage costs stack up across every duplicate. The fix for one problem becomes the seed of the next.
This is exactly the gap that AI-native storage architectures are built to close: keep one copy of the data, keep compute next to it, and avoid paying a toll every time someone wants to read from it. The impedance mismatch between file-producing systems and object-read workloads is precisely what a serverless filesystem like Archil addresses. It mounts existing S3 and GCS buckets as a native POSIX filesystem. Unstructured data never has to get forced into an object paradigm, and training code can read straight from the storage layer without SDK changes or reformatting.
Comparing major providers on the three metrics for AI workloads
Running the current field of AI storage options through sustained throughput, metadata latency, and co-location fit produces a clearer picture than any spec sheet comparison would.
Archil mounts S3, GCS, R2, Azure Blob, and S3-compatible buckets as a real POSIX filesystem. No migration step, no ETL job, no code changes. The bucket stays the source of truth the whole time. On sustained throughput, reads that hit the NVMe cache come back at sub-millisecond latency; a cache miss pulls from the source bucket and caches the result, and because the cache is elastic and billed only on what's actively used, capacity scales with the workload instead of requiring guesswork up front. On metadata latency, Archil supports full POSIX semantics, including atomic rename, flock/fcntl, hard links, and symlinks. Existing training code, checkpoint writers, and data loaders run unmodified. No rewrite, no new client library. On co-location fit, this is where the architecture does its most important work: cache nodes sit directly on the GPU cluster, and no data crosses regions without explicit configuration. Data residency stays inside the customer's account and can be revoked at any time, because the vendor never holds a persistent copy outside it. Writes get committed redundantly across Availability Zones on fsync(), then flushed asynchronously back to the bucket, so durability doesn't come at the cost of blocking the training loop.
Co-location is where architecture decisions matter most: if a storage layer can place cache nodes directly on the compute cluster while leaving the source bucket inside the customer's own account, it removes the egress path altogether, and a generic object store can't do that without copying data outside its home region. Archil is also storage-agnostic: S3, GCS, R2, and Azure Blob all mount the same way, which matters directly in multi-cloud environments where teams want one data layer spanning clouds without building replication pipelines to hold it together. Archil also ships serverless execution alongside the filesystem, so an agent or training job can run commands directly against its own filesystem without standing up a separate sandbox, and that takes on the agent-context and persistent-workspace problem that none of the other options in this comparison address directly.
CoreWeave AI Object Storage, expanded in an announcement on September 16, 2026, was built specifically for AI and co-designed with CoreWeave's compute and InfiniBand networking. Its Local Object Transport Accelerator (LOTA) runs as a proxy on every node in a CKS cluster and builds a local cache across node disks, delivering up to 7 GB/s per GPU, with throughput scaling linearly as workloads grow. As of an October 16, 2025 announcement, a single dataset can be reached from any CoreWeave region, other public clouds, and on-premises environments without replication overhead, and cross-regional LOTA acceleration reached general availability on October 16, 2026, with multi-cloud and hybrid support following in early 2026. CoreWeave reports that AI Object Storage has become the preferred storage layer for most of its large-scale model training workloads, alongside inference and agentic use. It carries full S3 compatibility and built-in observability through Prometheus and Grafana.
Google brought two offerings to Cloud Next 2026, and they sit at different points on the performance spectrum. Cloud Storage Rapid splits into Rapid Bucket, which cuts blocked GPU time and speeds up data loading for multi-modal training along with checkpoint restores, and Rapid Cache, which accelerates reads on existing buckets with no code changes and speeds up model loading for inference. Rapid Storage performance jumped from 6 TB/s to 15 TB/s at the Next 2026 announcement. Managed Lustre delivers up to 10 TB/s of throughput to TPU 8t and A5X environments over RDMA, Google's answer to the parallel-file-system tier needed right at the compute boundary, tightly integrated with its TPU and A5X environments so the RDMA path skips the CPU overhead of translating between object and file semantics. Thinking Machines Lab is a named early customer, and it confirms the Rapid Cache results when it runs real multi-modal training.
AWS S3 Express One Zone comes closest to a high-performance object tier on AWS, and it cuts per-object latency substantially versus S3 Standard. SageMaker jobs reading from Express One Zone-colocated storage see noticeably shorter training times from eliminated I/O wait, and Athena queries run faster against the same data. You need single-AZ deployment alongside compute to get that benefit. Storage costs carry a real premium over S3 Standard in us-east-1, and that premium is worth paying when reduced GPU idle time offsets it, which tends to hold for high-utilization training jobs and tends not to hold for archival or rarely accessed data.
DDN serves large GPU cluster deployments through two products built for different jobs: Infinia for inference and model loading, EXAScaler for large-scale training and checkpointing. Infinia provides multi-tenancy, metadata indexing, and sub-millisecond latency, with integrations into NeMo, NIM microservices, Trino, Apache Spark, TensorFlow, and PyTorch. Its split architecture directly answers the protocol bifurcation that practitioners at GigaOm describe: parallel file systems at the active compute boundary, object storage at the data lake tier.
Scality RING XP serves organizations that run on-premises or in a private cloud, where data residency or regulatory limits rule out public cloud object storage. Cloudian HyperStore is the most cited on-premises S3-compatible option in the reviewed material: it hits nearly 35 GiB/s per storage node on a 6-node cluster, and it serves teams that need full S3 compatibility on-premises and won't route training data through a public cloud.
If you keep the source bucket as the one authoritative copy, with no vendor holding a persistent replica outside the customer's account, the math on multi-cloud sprawl changes. The same dataset mounts from S3 or GCS without duplication, without egress fees between clouds, and without the synchronization and drift risk that comes from maintaining separate copies in separate places.
Where parallel file systems still beat object storage at the compute boundary
Object storage, even the fastest AI-tuned version of it, has a limit. At very high throughputs, any layer sitting between compute and object storage, translating one protocol into another, adds CPU overhead. At AI scale, that overhead can turn into the actual bottleneck rather than staying a cheap abstraction cost.
Practitioners at GigaOm describe this as a split running through every serious AI pipeline. Object storage dominates the ingestion and data lake tier, where exabyte-scale horizontal scaling and rich metadata for semantic discovery matter most. Parallel file systems stay necessary right at the active compute boundary, where random write, file locking, and low-latency metadata decide whether a training step finishes on time.
Google's Managed Lustre, delivering up to 10 TB/s over RDMA to TPU and A5X environments, is a clear example of this boundary in practice. No comparison in this piece is arguing that object storage replaces a parallel file system at the exact point where a GPU reads its next batch under the tightest latency budget in the pipeline. The honest shape of the landscape has object storage, including POSIX-mounted approaches like Archil, winning the ingestion and training-data layer where co-location and metadata latency decide GPU utilization, and parallel file systems holding their ground at the compute boundary where locking and random write still call the shots.


