Towards Looped Models Done Right: Rethinking at Fixed Points for Efficient Training, Decoding, and RL
循環語言模型的「定點」重塑:邁向高效訓練、解碼與強化學習的全新架構
Looped language models incur high costs per recurrence. This study leverages the insight that as recurrent states approach fixed points, the exact path taken matters less. To exploit this, the authors introduce a learned depth prior (via prediction feedback with entropy regularization) and orthogonal input injection. Tested across 100M to 1.6B parameters, these techniques consistently lower perplexity. They enable 3x smaller KV cache via terminal sharing, 1.79x faster student prefill, and 2x faster RL gradient computation directly from saved rollout states.
Key points
Fixed-Point Simplification
As recurrent states near fixed points, the trajectory matters less, enabling truncated backpropagation and terminal KV sharing with near-zero accuracy loss.
Feedback-Learned Depth Prior
Replaces fixed-depth and overly broad priors with a learned prior guided by prediction feedback, utilizing entropy regularization to optimize supervision at the target depth.
Orthogonal Input Injection
Traditional injection schemes suffer from state components amplifying or canceling the input. Orthogonal injection removes this component for stable recurrence.
Accelerated RL & Prefill
RL updates compute gradients directly from saved rollout states (2x faster), while a distilled student model speeds up prefill by up to 1.79x.
Why it matters
This work directly addresses critical bottlenecks of looped models, including KV cache explosion, slow prefill, and expensive RL backpropagation. Matching full fixed-depth performance at 1.6B scale with a 3x smaller KV cache demonstrates that memory-efficient looped models are highly viable for consumer hardware and high-throughput production environments.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
How to use it
- 1Edge AI deployment: Run 1.6B-scale looped LLMs with 3x smaller KV cache footprint on memory-constrained devices.
- 2Low-latency chatbots: Speed up the prefill stage by up to 1.79x using a distilled student, reducing Time-to-First-Token.
- 3Efficient RLHF alignment: Compute gradients directly from saved rollout states during RL training to double the training throughput.
Limitations & caveats
- The fixed-point convergence and prior learning mechanics increase model complexity, requiring careful hyperparameter tuning during training.
- The evaluation is limited up to 1.6B parameters; scalability and stability on ultra-large scale models (e.g., 70B+) require further validation.
Related

Falcon-Emirati-7B: Bridging the Gap in Emirati Arabic Dialect and Culture
解鎖阿聯酋方言與文化:專為在地語境打造的 Falcon-Emirati-7B 模型
Falcon-Emirati-7B is a 7B parameter model specialized in Emirati Arabic, capturing local dialect, Nabati poetry, and cultural nuances where generic models fail.

Google Launches EmbeddingGemma 2: Compact 740M Parameter Multimodal Embedding Model for Ultra-Low Latency Edge AI
Google 推出 EmbeddingGemma 2:超輕量 740M 參數,讓裝置端擁有強大「多模態語意搜尋」與即時決策力
Google's EmbeddingGemma 2 is a 740M open-weight multimodal embedding model that runs entirely on-device, unifying text, image, video, and audio into a single vector space with minimal memory footprint.
Base Models Can Reason: Unlocking Latent Performance with Strategic Starting Tokens
基礎模型也能推理:啟動關鍵「開頭 token」釋放隱藏實力
A new study reveals that forcing base models to start with specific token cues like 'Okay' triggers reasoning behavior comparable to RL-tuned models, tracing this effect directly to pre-training data structures.