Long sessions, scarce memory
A long-running agent works with 100,000 tokens of context or more, and the GPU memory that holds that context is in short supply. Our new paper keeps each session's KV cache in host memory and lends it GPU memory while there is room.
The fast memory stacked next to every GPU, HBM (high-bandwidth memory), is in short supply, and the next generation of hardware shows that this is not about to change. NVIDIA’s Rubin Ultra will carry 192 GB of HBM per GPU, less than first planned and less than the 288 GB of today’s B300.
Model designers are already planning around it. DeepSeek’s Engram lets about a quarter of DeepSeek-V4.1-Flash’s weights live in CPU memory, or even on SSD, instead of HBM.
The other large consumer of HBM is the KV cache: the keys and values a model computes for each token of a session, kept in GPU memory and reused at every step to produce the next token. It grows with every token, and agents run long sessions. That is where our new paper, Adaptive DHiSparse: Load-Adaptive KV Management for Sparse and Speculative Serving, comes in.
Our approach keeps each session’s complete KV cache in host memory, the server’s CPU memory, lends it GPU memory while there is room, and takes that memory back when other sessions need it. We developed it into a complete component of a modern serving stack: it works with the latest speculative decoding drafters, DSpark and DFlash 2, and with disaggregated serving, where prompts and answers run on separate GPUs.
As GPUs get less memory for their compute and agent sessions keep growing, keeping every KV cache in GPU memory leaves more and more of that compute idle. Putting that compute to use gets more tokens out of every GPU, and with demand for tokens growing faster than the supply of compute, that is how serving keeps up. We expect this kind of memory management to become a standard part of long-context serving.
We start with what long sessions ask of GPU memory, then walk through how sparse attention lets most of a session’s history live off the GPU, what HiSparse already does with that, what we changed, and what we measured.
1. Long sessions need a lot of GPU memory
A language model writes one token at a time, and each new token depends on everything that came before it, so the KV cache grows with the session, one entry per token. It sits in HBM, the GPU’s own memory, which is very fast, expensive and limited, and the model’s weights take their share of it first.
Agents make sessions long. A coding agent rereads its whole workspace before each step, the files it opened, the outputs of its tools and its own earlier reasoning, and then writes a few hundred tokens. On our platform, an average request carries about 134,000 tokens of context for some 800 tokens of output, as we described in Tokenomics, behind the scenes. On GLM-5.3, a single 100,000-token session needs about 4.4 GiB of KV cache.
A GPU serves many sessions at once. At each step, it reads the model’s weights once and applies them to every session in its batch, so the more sessions it advances together, the more tokens it produces for the same work. That only holds while every session’s KV cache fits in memory. Once memory is full, new requests wait in line, and the GPU’s output stops growing even though its compute could do more.
Figure 1 of the paper shows this on GLM-5.3 with 32,000-token prompts. Resident serving, the conventional setup that keeps every session’s KV cache in GPU memory, grows with the load until about 64 simultaneous requests. Doubling the requests to 128 then adds less than 5% more output, while the average wait for a first token goes from under a second to over a minute, and DFlash 2 speculative decoding does not move that limit. The same GPUs keep growing with our memory system: at 128 requests, Adaptive HiSparse, which uses it alone, delivers 1.9 times the output of resident serving, and Adaptive DHiSparse, which adds speculative decoding, 1.6 times that of resident serving with DFlash 2.
2. Sparse attention and HiSparse
What makes that possible is a change in how recent models read their KV cache. They use sparse attention: a small, fast component called the indexer scores every earlier position and picks the ones that matter most for the next token, and attention reads only those. GLM-5.3 picks 2,048 positions per step. DeepSeek-V4.1-Flash picks 512, about 0.1% of a 500,000-token session.
The rest still has to stay within reach, since the indexer may pick any position at the next step and the picks change every time. But it does not have to sit in HBM. Host memory, the server’s CPU memory, is a second pool next to it: on a B300 server, it is of the same order of size as the HBM of all the GPUs combined, a few terabytes. It is cheaper per gigabyte and slower for the GPU to reach, which suits the part of the history the model is not reading right now. HiSparse (Xie et al., 2026) keeps every session’s complete KV cache there. On the GPU, each session keeps the compact index the indexer searches, so the model still considers its whole history, and a small hot cache of recently used entries. At each step, picked entries already in the hot cache are used in place, and the missing ones are copied from host memory just before attention needs them. The model computes exactly what it would have computed with everything on the GPU; only the location of the bytes changes.
Offloading the history to host memory frees most of a session’s GPU memory. On GLM-5.3 at 100,000 tokens, a session needs 4.44 GiB on the GPU when its whole KV cache stays there, and 0.46 GiB with HiSparse, about ten times less.
HiSparse leaves one decision to whoever deploys it: the size of the hot cache, the same for every session and every load. A small cache fits more sessions, but sends more reads to host memory, even at quiet times when the GPU has memory to spare. A large cache saves those reads, but fits fewer sessions when the server is busy. In our paper’s words, the cache size is “a deployment-level compromise rather than a direct response to the current memory demand”.
3. Adaptive DHiSparse: lending the free GPU memory
The paper rests on one split of responsibilities: the model decides which entries to read, and the serving system decides where those entries are kept. Adaptive DHiSparse never changes the first. It changes the second continuously, as the load changes.
Each session keeps a working set that is never lent: its hot cache and the newest tokens it is writing. Whatever GPU memory is left over holds extra copies of the sessions’ history, which we call borrowed history. When the server is quiet, a session can borrow room for most or all of its history, so most of the entries the indexer picks are already on the GPU: in a load test on GLM-5.3, lending cut the share of picks fetched from host memory at low load from 14.2% to 1.4%. When only part of the history fits, one pass on the GPU sorts the picks out in order: the newest tokens and borrowed history first, then the hot cache, and host memory only for what is still missing. That search stays as small as the hot cache, however much history a session has borrowed.
Borrowed history is given back when it is needed. When a new session needs room, the running sessions return pages, each in proportion to its history. Nothing is lost: a page is backed up to host memory before it is returned, and if it is picked later, it is copied back like any other miss. When a session ends, its space goes to the others.
The same load test shows the loan being recalled: when requests jumped from 16 to 256, the share of history held on the GPU fell from between 75 and 100% to under 1% within 16 seconds, and climbed back above 50% a little over two minutes after the load dropped.
As a result, the choice HiSparse leaves to deployment goes away. The hot cache can stay at the small size that suits a busy server, and at quiet times, borrowed history fills the memory a larger cache would have taken. At low load, sessions run almost as if their whole KV cache were in GPU memory. At high load, they shrink toward the HiSparse working set, and the server fits more of them. That is what we were after: a memory system an inference stack can leave on without tuning it for the traffic it expects.
4. Speculative decoding
Many inference stacks also use speculative decoding: a small drafter proposes several tokens, and the model checks them all in one pass, which yields more tokens per step whenever the guesses are right. Our runs pair DeepSeek-V4.1-Flash with DSpark and GLM-5.3 with DFlash 2. The HiSparse version we compared against has no path for a drafter: it keeps no KV cache for one and looks up a single query per session at a time. Adaptive DHiSparse brings draft-model speculative decoding to this kind of serving. Speculation makes memory harder in two ways. Each drafted token picks its own KV entries, so one check reads more of the history. And the drafter’s own KV cache, plus the scratch space for the checks, takes memory that could hold another session. A memory system that stays on has to handle both.
One copy per missing entry. When several drafts pick the same entry and it is not on the GPU, it is copied from host memory once, and each draft still attends to exactly its own picks. In our counters, this avoided 29% of the copies from host memory during checks.
A bounded drafter cache. The drafter’s KV cache lives in a ring sized to what the drafter can see. GLM-5.3’s drafter looks back 2,048 tokens, so its ring holds 2,176 rows per session, however long the session grows.
Reused scratch space. The space that holds picks during a check is handed from one group of layers to the next once its readers finish. On GLM-5.3, four such areas replace 78, which cuts that part of the memory by 95%.
Rejected drafts never enter the history, and neither drafter changes: DSpark and DFlash 2 propose and accept tokens exactly as they were designed to.
5. What we measured
We measured both models with long prompts whose prefix was already cached, from a few simultaneous requests to hundreds, and compared Adaptive HiSparse and Adaptive DHiSparse with resident serving, where every KV cache stays in GPU memory, and with HiSparse. The GLM-5.3 runs used disaggregated serving, with separate GPUs for prompts and answers.
On GLM-5.3 with 100,000-token prompts and DFlash 2, resident serving stops growing at about 16 simultaneous requests, near 220 tokens per second per GPU, and at 128 requests, the average wait for a first token reached about four minutes. Adaptive DHiSparse delivered 793 tokens per second per GPU at 128 requests, with a 5.1-second average wait, and 826 at 256.
On DeepSeek-V4.1-Flash with 500,000-token prompts and 512 simultaneous requests, Adaptive DHiSparse delivered 5,054 tokens per second per decode GPU, against 2,738 for resident serving with the same DSpark speculative decoding. Without speculative decoding, Adaptive HiSparse delivered 4,217 against 1,853.
The gain grows with the context. On DeepSeek at 512 requests without speculation, with 32,000-token prompts, where memory is not yet the limit, the two serve about the same: 6,320 and 6,489 tokens per second per decode GPU. At 200,000 tokens, Adaptive HiSparse serves 57% more.
Against HiSparse, which cannot use speculative decoding, the difference is larger. On GLM-5.3 at 128 requests, Adaptive DHiSparse delivered 45% more, 793 against 548 tokens per second per GPU. Under the same latency limits for every configuration (99% of first tokens within 10 seconds once the initial burst has cleared, and 40 milliseconds per token on average), it handled four times as many simultaneous requests as HiSparse, with 3.7 times the output.
Without speculative decoding, Adaptive HiSparse delivered 19 to 23% more than HiSparse at light load and 13 to 14% more at 128 and 256 requests on GLM-5.3, and 9% more on DeepSeek at 512 requests, where the first token came after 20 seconds instead of 28. Part of the GLM gap comes from a different attention kernel in our runs, so we do not credit all of it to the memory policy. A well-tuned fixed cache can still match it at some points: at 200,000 tokens on DeepSeek, the fixed cache came out 0.3% ahead.
6. What it costs
Serving more sessions at once means each one moves a little slower, because every step is shared by more of them.
At 512 requests on DeepSeek with DSpark, each session gets a new token every 24 milliseconds instead of every 12.5, and requests wait 25 seconds instead of 80 for their first token. Adaptive DHiSparse does not make a single session faster. It lets the same GPUs serve more sessions, and the right balance depends on the workload.
At light load, there is a small price. Without speculative decoding, adaptive serving was 3 to 7% slower than resident serving while every KV cache still fit: 3.6% on DeepSeek at 16 requests, about 5 to 7% on GLM-5.3. With DFlash 2 on GLM-5.3, the gap was about 9 to 10% at 4 and 8 requests, while with DSpark on DeepSeek, adaptive serving came out ahead even at 16.
Speculation has its own trade-off. Its extra memory leaves room for 140 GLM-5.3 sessions instead of 208. At 256 requests, Adaptive DHiSparse and Adaptive HiSparse delivered about the same, 826 and 823 tokens per second per GPU, within our run-to-run spread, but requests waited 65 seconds for their first token instead of 34, almost all of it queuing for one of the 140 slots. We also picked DFlash 2’s block size by hand for each load (8 up to 32 requests, 6 at 64, 4 at 128 and 256).
Our runs cover one deployment per model, with prompts whose prefix was already cached. What does not change is the computation: every step attends to the same entries, with the same bytes, as it would with everything on the GPU. The paper states the exact conditions.
7. What comes next
In long-context serving, the same GPU memory can hold more sessions, larger hot caches, more borrowed history or the drafter that speculative decoding adds, and the mix that gets the most output changes with the load. Adaptive DHiSparse adjusts one part of that mix as the load changes: how much history each session keeps on the GPU. The rest, such as the size of the hot cache and the drafter’s settings, is fixed for the whole deployment, and finding the best mix at each load is still an open question. That is what we want to work on next, along with running the system on real agent traffic. The same approach may also let local machines run longer sessions, since the system RAM next to their GPU is often larger than the GPU’s own memory, though we have not tested it there.
The full paper is coming to arXiv soon, with the mechanism, the correctness argument and the full results. We will link it here as soon as it is live.
Sources
- Rubin Ultra and HBM per GPU: SemiAnalysis, Long live the short king: why 4-hi HBM wins, September 2026.
- Engram: Cheng et al., Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models, 2026.
- HiSparse: Xie et al., HiSparse: Scaling sparse-attention decoding with hierarchical KV cache management, 2026.
- The models: DeepSeek-AI, DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression, 2026; Z.ai, GLM-5.3, 2026.
- Speculative decoding: Cheng et al., DSpark: Confidence-scheduled speculative decoding with semi-autoregressive generation, 2026; Inco AI, DFlash 2: Keep drafting parallel, 2026.
- Our earlier posts: The token gap, on why serving has to get more tokens from every GPU; Tokenomics, behind the scenes, for our production request mix; and Eleven models in eight weeks, for the KV cache size of each model we serve.