NVMe Caching Layers in Cloud Filesystem Architectures
NVMe caching bridges the latency gap between object storage and AI's random-access demands.

Object storage was built for one job, and AI training is asking it to do a different one. S3 and its cousins are great at being cheap, durable, and able to grow forever without anyone having to think about it. What they're not built for is the way a training loop or an inference engine actually touches data: constant, small, random requests, fired off thousands of times a second, each one expecting an answer right now. Object storage was tuned for big sequential reads, the digital equivalent of reading a book front to back. AI workloads read like someone flipping to a random page every few milliseconds, over and over, for hours.
The gap between those two worlds isn't small. NVMe drives answer in tens to hundreds of microseconds. That's not a rounding error, that's orders of magnitude, the kind of gap that separates a car from a jet.
And the usual engineering trick of throwing more hardware at a slow system doesn't help here. You can parallelize throughput by adding threads, but you can't parallelize away latency. Each individual round trip to object storage still takes as long as it takes, no matter how many of them you run at once. Stacking ten thousand slow phone calls on top of each other still leaves any single call just as slow.
The problem goes deeper than speed, too. S3 and similar systems were never built with the file semantics that training frameworks expect. There's no atomic rename, no file locking, none of the POSIX behavior a filesystem is supposed to guarantee. Training code and AI agents assume they're talking to a filesystem. Object storage is something else wearing a filesystem's clothes.
The real cost appears on the GPU side, in idle cycles paid for while waiting on storage. A model with trillions of parameters needs its weights pulled from storage into GPU memory before a single gradient step can happen. GPUs are some of the most expensive compute on the planet, and when storage can't keep up, those GPUs sit there waiting. Slow storage doesn't just slow down a job, it burns expensive silicon doing nothing.
What a caching layer does in the stack
The fix isn't to replace object storage. It's to put something faster in front of it. An NVMe caching layer sits between compute and the object store, catching most of the traffic before it ever has to make that slow round trip to the bucket. The bucket doesn't change. Nothing about its durability, its cost, or its scale gets touched. It just stops being the thing compute has to wait on for every read.
A cloud filesystem usually mounts right over an existing object-storage bucket, and a local or distributed NVMe tier sits in front, intercepting reads before they reach it. Writes work in a similar spirit. Data gets copied to more than one place for safety, the write is confirmed back to the application, and then it gets flushed to the bucket later on its own schedule. The bucket stays the one true, durable copy of everything, the record that never gets second-guessed.
This setup keeps the economics of object storage, cheap, durable, able to grow without limit, while giving compute something closer to file-system speed to work with. Neither tier replaces the other. They work together, each one doing the part it's actually good at.
From the outside, a cloud filesystem usually looks simple: it shows up through NFS or SMB, or sometimes a vendor's own client, and looks like an ordinary mounted drive. Architecturally, these systems tend to break into five layers: the client and its access protocol, the namespace and metadata system that tracks what's where, the caching tier, the data servers or storage nodes, and the durable backend. The NVMe cache lives in that middle caching tier, acting as the traffic cop between what the client is asking for and what the backend can actually deliver.
The three-tier storage hierarchy that practitioners now deploy
In practice, teams running AI infrastructure have settled on a fairly consistent three-tier setup. NVMe holds the hot working set, the data a job needs right now. SATA or warm SSD holds data that's been staged for use soon but doesn't need microsecond access. Object storage holds the durable archive, everything that needs to exist forever but doesn't need to be touched often. The caching layer is what actually moves data between these tiers as access patterns shift.
That movement follows a pretty simple logic: how often is this data touched, and how much delay can the job tolerate waiting for it? Raw archives, full copies of datasets, and old checkpoints that might never be touched again belong in cold object storage, managed by lifecycle automation.
The single biggest lever for controlling cost in this setup is keeping the hot tier small and honest. Only the data a job needs right now should sit on NVMe. Everything else should fall down to cheaper tiers automatically, without anyone having to manually sort it. Teams tend to build toward one of four recipes depending on what they're optimizing for. A hybrid cloud setup uses disaggregated NVMe to absorb bursts of demand, backed by object storage underneath for durability and the ability to serve data across regions.
Checkpoint storage is a good small-scale example of how all this tiering logic plays out. None of this requires someone sitting there deciding what to move and when. Lifecycle policy handles it automatically, the same way a thermostat handles temperature.
How design choices inside the caching layer govern performance
Putting NVMe into the architecture is the easy part. Whether it actually delivers the speed it promises comes down to a set of design choices inside the caching layer itself, how data gets placed, how writes are handled, how the cache talks to everything around it. Get these wrong and the fast drives underneath might as well not be there.
Write behavior matters just as much as read speed. A better approach copies the write to a couple of places for safety, confirms it back to the application immediately, and flushes it to the bucket afterward on its own schedule. That balances the need for durability against the need for speed.
NVMe drives are so fast that the bottleneck in these systems often moves off storage entirely and onto everything around it: CPU, memory, and network. Every single read or write still needs CPU cycles to handle data placement, erasure coding, checksums, and metadata bookkeeping. Without the right tuning, that overhead can eat up most of the performance the drives are theoretically capable of, leaving a fast drive doing a slow drive's job. Scaling performance up usually means scaling CPU threads up right alongside the number of NVMe drives, not just adding more drives and hoping the rest of the system keeps pace.
Reads and writes also aren't equally fast, and that's baked into how replicated NVMe systems work. With three-way replication, a write has to be confirmed across all three copies before it counts as done, which makes writes structurally slower than reads. Random reads can run about five times faster than random writes on the same system, and large sequential reads can be nearly three times faster than sequential writes. That asymmetry decides which workloads actually benefit most from this kind of caching: read-heavy jobs like serving training data see outsized gains, while write-heavy jobs like constant checkpointing see a smaller bump.
Metadata is its own separate performance problem, distinct from how fast raw data moves. Listing a directory with millions of files, or creating thousands of small files in a hurry, can choke a system's metadata service even when the actual data volume involved is tiny. Systems that separate metadata handling from data handling, letting each scale on its own, tend to hold up much better under these metadata-heavy AI workloads than systems where the two are bundled together.
Local NVMe versus distributed NVMe cache: the central trade-off
The single biggest architectural fork in the road is whether to cache data on NVMe local to each compute node, or in a shared, distributed NVMe tier. Neither option is right in every case, and the trade-offs run deep enough that this decision shapes almost everything else about how the system behaves.
Local, per-node NVMe has some clear strengths. A cache hit never has to cross the network, so latency is bound only by how fast the drive itself can respond. Checkpoint workloads, which write heavily, benefit from not having to pay a replication tax on every write. The consistency model stays simple too, since there's no need for a distributed protocol to keep multiple caches agreeing with each other. For jobs that live entirely on one machine, like single-GPU training or local inference, local NVMe delivers its full benefit with no coordination overhead.
A shared, distributed cache solves a different problem. In an elastic cloud environment, every new compute node starts with an empty cache and has to warm it up from scratch, a cold-start cost that repeats every single time the cluster scales out. NVMe-oF lets storage be attached and shared across nodes, and paired with fast fabrics like RoCE or InfiniBand, multiple nodes can reach the same dataset without each one needing its own full local copy. Inside a single cloud datacenter, the network latency between nodes can be low enough that fetching from a shared NVMe tier over that network still beats a cold fetch from object storage by a wide margin.
There's a durability catch that applies no matter which path gets chosen. NVMe attached to a cloud VM is ephemeral. It doesn't survive the VM being deallocated or stopped. That means the NVMe tier, local or distributed, solves latency, while durability stays with object storage. Object storage has to remain the authoritative, durable copy of the data regardless of which caching approach sits on top of it.
The practical answer splits along workload stability. Local NVMe wins for stable, single-node or fixed-cluster workloads, where the cost of warming up the cache gets paid once and amortized over a long run. Distributed NVMe wins for elastic, multi-node workloads, where nodes come and go and no single machine can be counted on to hold the data. Either way, the NVMe tier has to be treated as a cache and never mistaken for permanent storage.
KV cache offloading: inference opens a second front for NVMe
Training isn't the only place NVMe caching earns its keep. Inference creates its own demand for it, and the driver isn't dataset size at all, it's the memory footprint of the key-value cache that large language models build up as they hold long conversations or process long documents.
For a model handling long context windows, a single user's KV cache can eat up tens of gigabytes of GPU memory on its own.
The fix follows the same tiering instinct seen in training. The hottest KV blocks stay on GPU HBM, where access is fastest. Warm blocks move down to CPU DRAM. Cold blocks, older context, prefix caches that might get reused later, move down to NVMe SSD. The same hot-warm-cold pattern used for training data applies here to a model's working memory.
DDN's work with NVIDIA Dynamo is a concrete example of this in action. It extends a GPU's KV cache out into NVMe-based storage through a four-tier hierarchy: HBM, then DRAM, then local SSD or networking-accelerator-attached shared NVMe SSD storage, then external NVMe SSD storage further out. That makes NVMe-resident KV cache part of the model's addressable context memory, not just a backup copy sitting off to the side, and it persists across inference runs on DDN's Infinia and EXAScaler platforms.
That persistence is where things get genuinely interesting. When KV blocks stored on NVMe survive past the end of one session, a returning user or a workload that reuses the same prefixes repeatedly skips having to recompute all that context from scratch. At that point, the NVMe tier stops just being a cache for one conversation and starts acting like a form of long-term memory for an AI agent, something that remembers rather than something that merely holds a session's leftovers.
What production deployments show about the caching layer's behavior
The theory holds up in production, with one consistent asterisk: everything works as described once the cache is warm, and the warming process itself carries a cost worth watching.
CoreWeave's AI Object Storage, built around something it calls the Local Object Transport Accelerator (LOTA), caches frequently used or pre-staged data on the local NVMe disks sitting inside GPU nodes. That number describes a fully warm cache. It says nothing about the first request for data that hasn't been cached yet, which still has to make the slower trip back to object storage.
CI build infrastructure tells a similar story from a completely different industry. A cache sitting in a registry or an S3 bucket costs a network round trip, every single build, whether the code changed or not.
The CI example points to a rule that applies everywhere NVMe caching is used: the benefit only compounds if the cache outlives the compute instance using it. A machine that gets torn down and rebuilt from scratch for every job throws away a warm cache before it ever gets the chance to pay for itself. The caching layer decides whether fast storage actually gets used, so a job has to show up and use it or that storage just sits there, fast and empty.
Sources
- In-kernel caching for distributed cache
- Getting the MOST out of your Storage Hierarchy with Mirror-Optimized Storage Tiering
- Hybrid storage system
- Caching and tiering for cloud storage
- Assise: Performance and Availability via NVM Colocation in a Distributed File System
- DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
- Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs
- LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference


