DeepSeek Shrank the Cache. Here's Why That Matters More Than the Benchmarks
Most people judge a language model by its benchmark scores. That's fair. Scores are what you see, and they're easy to compare.
But if you've ever paid for inference at scale, you know the number that actually hurts lives somewhere else. It's memory. More specifically, it's the KV cache, the pile of stored attention data a model drags around so it doesn't have to reread everything from scratch.
DeepSeek's new technical report on V4.1-Flash, released in September, is mostly about that one problem. I think it's one of the more interesting things they've put out, and not because of the headline score.
Agents broke the old math
A normal chat is short. You ask, the model answers, and that's mostly it.
Agents don't work like that. A coding agent reads a file, runs a command, reads the output, edits something, runs the tests, and loops for an hour. Every step pulls the whole history back in. The model spends most of its time reading, not writing. DeepSeek calls these workloads input-heavy, and that's the right word.
Here's why that matters. Every token the model has already seen leaves behind keys and values, stored so attention doesn't have to recompute them. That's the KV cache. While a request is running, it sits in HBM, the fast memory on the GPU, which is small and expensive. When a session pauses, the cache gets pushed out to host memory or SSD so it can be picked up again later.
So you hit two walls. HBM caps how much history you can serve at once. Disk caps how much history you can keep around to reuse. And moving all of it back and forth eats bandwidth.
DeepSeek's argument is simple. Sparse attention already made long context cheap in terms of raw compute. What's left is storage and data movement, and that's now the main thing standing between us and cheaper agents.
They've been at this for a while
None of this comes out of nowhere. DeepSeek has been obsessed with memory and efficiency for a couple of years now.
V2, back in 2024, introduced Multi-head Latent Attention. Instead of storing full keys and values for every head, it squeezed them into a small latent vector. That alone cut the cache down a lot. The same model used a fine-grained Mixture-of-Experts setup, where each token only wakes up a small slice of the network.
V3 kept going. It balanced its experts without the usual auxiliary loss, and it trained with multi-token prediction, which teaches the model to look a few tokens ahead and can also speed up generation. Around the same time they open-sourced a bunch of their low-level kernels, the boring GPU code that decides whether your hardware is actually busy or just waiting on memory.
V4 moved to sparse attention. V4.1 is where they turn the screws on the cache itself. By their own chart, the global cache per token is now roughly 437 times smaller than it was in the very first DeepSeek model.
Four tricks, one goal
V4.1-Flash is a 552B-parameter multimodal MoE with a one million token context window, pretrained on 45 trillion tokens. That's the spec sheet. The interesting part is what they did to the cache, and it comes down to four ideas.
1. Split the model in two
They call it a Causal Encoder-Decoder. The 40-layer backbone is cut in half: 20 layers act as an encoder, 20 as a decoder. The decoder doesn't build its own global cache from scratch. It gets it projected from the encoder's final hidden states.
The payoff is in the active parameter counts. Reading the prompt (prefill) uses 8B active parameters per token. Generating (decode) uses 16B. For an agent that reads far more than it writes, halving the cost of reading is a big deal. One honest note: these are parameter counts, not measured wall-clock speedups. The paper doesn't claim prefill is literally twice as fast.
2. Stop storing the same thing in every layer
The second idea is Compressed Sparse Attention 2, or CSA2. In a normal model, every layer keeps its own copy of the cache. A lot of that is redundant.
CSA2 lets each layer run in one of three modes:
- Full: the layer builds its own compressed cache and picks which past tokens to attend to.
- Reindex: it borrows the cache from an earlier layer, but uses its own query to pick different tokens.
- Reuse: it borrows both the cache and the earlier layer's picks.
Every layer still computes its own query and its own local window cache. Think of a study group where one person takes notes, some people highlight their own parts of those notes, and the rest just read the highlights.
3. Use fewer bits
The main cache is stored in FP4, four bits per number. The short-range sliding window cache stays at FP8. Put that together with CSA2 and the global cache lands at 890 bytes per token, about a quarter of what V4-Flash needed.
By my rough math, a full million-token context works out to under a gigabyte of global cache. That's the kind of number that changes how many sessions one GPU can hold.
4. Throw away what you can rebuild
This one is my favorite, because it's an engineering call rather than a modeling trick.
In V4, the sliding window cache (SWA KV) made up almost half of the persistent storage. But that data only matters for the most recent stretch of tokens. Long-lived reuse comes from the global cache, not the local one.
So V4.1 stops saving SWA KV to disk. It keeps it in a shared pool of host memory for a few minutes, then lets it expire. If you come back to a session and the global cache is still there but the local one is gone, the model rebuilds it with what they call SWA Bounded Replay. It only recomputes the last window of tokens, instead of the much larger amount an exact rebuild would need.
The catch is that the rebuilt state is approximate. It's not bit-for-bit what you'd have gotten originally. DeepSeek says the quality hit is barely noticeable. The win is that persistent storage drops to about an eighth of V4-Flash's. You're trading a bit of recompute for a lot of disk.
So did it cost them anything?
The obvious worry with all this compression is that the model gets dumber. According to their numbers, it didn't.
On DeepSWE v1.1, a software engineering benchmark run at maximum reasoning effort, V4.1-Flash resolves 74.2. V4-Flash scored 54.4. Opus-5 sits at 74.0. A model with a quarter of the cache beat its predecessor by about 20 points and roughly tied a top closed model.
I'd be careful with how far you take that, though. The architecture changed, the training data changed, and the post-training changed. There's no clean experiment that isolates the cache tricks and shows what each one cost or gained. So the right reading is "they compressed the cache a lot and the model still got better," not "compression is free."
DeepSeek is fairly upfront about the limits too. They point out that no test suite covers every case. If CSA2 picks the wrong tokens, or the approximate replay drifts, quality could slip in situations nobody has tested yet, especially right where a paused session gets picked back up.
Why I care about the plumbing
Benchmarks tell you how smart a model is. Cache size tells you how many people can afford to use it for real work.
If an agent session costs a quarter of the memory to keep alive and an eighth of the storage to park, that shows up in the price of an agent-hour. It decides whether a small team can leave an agent running on a big codebase all afternoon, or has to babysit the bill.
That's the part of AI that rarely gets a viral thread. It's also the part DeepSeek keeps winning. They've spent three generations treating memory as the enemy, and V4.1-Flash is the clearest version of that bet so far.
Sources
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression, DeepSeek-AI, September 2026