KV-streams: Boosting Context Compaction Efficiency in Agentic Reinforcement Learning
KV-streams:透過快取串流突破代理型強化學習的上下文壓縮瓶頸
Training agentic LLMs with long contexts requires context compaction to manage GPU memory. However, traditional compaction methods frequently flush and rebuild the KV cache (prefilling), severely slowing down training. To solve this, researchers introduced KV-streams. Instead of flushing the KV cache after compaction, it streams the cache forward. Compatible with multiple compaction strategies, KV-streams delivers a 2.6x to 5x training speedup. Intriguingly, the streamed KV cache functions as a recurrent state, carrying forward long-term memories lost from the active context—a capability that emerges naturally through RL alone.
Key points
Eliminating Prefill Bottlenecks
Traditional compaction methods require repetitive prefilling after each compaction. KV-streams avoids this by continuously streaming the KV cache forward.
Significant Training Speedups
Demonstrates 2.6x to 5x wall-clock speedups during training across three compaction strategies without compromising performance.
Emergent Recurrent Memory
The streamed KV cache acts as a recurrent state, retaining long-range memory eliminated from the active context, emerging naturally through RL alone.
Plug-and-Play Compatibility
Designed as a lightweight, plug-and-play enhancement compatible with various context compaction and post-training pipelines.
How it works
| 傳統壓縮方法 (Traditional) | KV-streams 方案 | |
|---|---|---|
| Cache Handling | 壓縮後立即清除並重新計算 | 持續向前串流(Streamed forward) |
| Prefill Cost | 極高,需重複進行 Context 預填 | 極低,避免重複運算 |
| Training Speedup | 1x (基準) | 達到 2.6x 至 5x 加速 |
| Long-range Memory | 受限於壓縮視窗邊界 | RL 可激發出循環狀態以保留超限資訊 |
Why it matters
Agentic LLMs are constrained by GPU memory during long-horizon tasks. While context compaction limits memory usage, it introduces severe training inefficiencies. KV-streams resolves this tradeoff, making RL training on long traces computationally viable. Crucially, proving that RL alone can foster recurrent-like memory retention in streamed KV caches opens new pathways for developing highly efficient, long-memory autonomous agents.
Who it affects
- AI Researcher
- AI Developer
- Enterprise Leader
How to use it
- 1RL training for long-horizon autonomous agents
- 2Reducing compute and memory overhead in multi-step reasoning tasks
Limitations & caveats
- The generalization of the emergent recurrent memory in highly complex, uncontrolled real-world environments remains to be fully verified.
- While highly compatible, the exact speedup may vary depending on the underlying hardware architecture and self-attention implementations.
Related

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude
艾倫人工智慧研究所開源 AstaBrief:比 Claude 快 3.5 倍的 8B 科學報告生成模型
Allen Institute for AI (Ai2) has open-sourced AstaBrief 8B, a specialized model for scientific report generation that achieves a 3.5x speedup over proprietary pipelines while maintaining high citation accuracy.
GALA: Distilling 3D Gaussian Avatars into Linear Blendshapes for Real-Time Animation
GALA:用線性混合變形蒸餾技術實現 3D Gaussian 虛擬化身即時動畫
GALA distills complex neural decoding of 3D Gaussian avatars into lightweight linear blendshapes, reducing CPU animation costs by up to 1000x and enabling 60fps real-time performance on mobile devices.
ScholarCatalyst: A Benchmark for Testing AI's Intuition in Retrieving Inspiring Research Papers
ScholarCatalyst:評估 AI 是否擁有「科學家直覺」的學術文獻檢索基準
ScholarCatalyst is a novel benchmark featuring annotations from 184 lead authors to evaluate whether AI can retrieve key inspiring papers from past literature based only on an initial research question.