Aivora
arXivLLMAdvanced

Towards Looped Models Done Right: Rethinking at Fixed Points for Efficient Training, Decoding, and RL

循環語言模型的「定點」重塑:邁向高效訓練、解碼與強化學習的全新架構

2 min read
Towards Looped Models Done Right: Rethinking at Fixed Points for Efficient Training, Decoding, and RL
The 30-second version

Looped language models incur high costs per recurrence. This study leverages the insight that as recurrent states approach fixed points, the exact path taken matters less. To exploit this, the authors introduce a learned depth prior (via prediction feedback with entropy regularization) and orthogonal input injection. Tested across 100M to 1.6B parameters, these techniques consistently lower perplexity. They enable 3x smaller KV cache via terminal sharing, 1.79x faster student prefill, and 2x faster RL gradient computation directly from saved rollout states.

Key points

01

Fixed-Point Simplification

As recurrent states near fixed points, the trajectory matters less, enabling truncated backpropagation and terminal KV sharing with near-zero accuracy loss.

02

Feedback-Learned Depth Prior

Replaces fixed-depth and overly broad priors with a learned prior guided by prediction feedback, utilizing entropy regularization to optimize supervision at the target depth.

03

Orthogonal Input Injection

Traditional injection schemes suffer from state components amplifying or canceling the input. Orthogonal injection removes this component for stable recurrence.

04

Accelerated RL & Prefill

RL updates compute gradients directly from saved rollout states (2x faster), while a distilled student model speeds up prefill by up to 1.79x.

Why it matters

This work directly addresses critical bottlenecks of looped models, including KV cache explosion, slow prefill, and expensive RL backpropagation. Matching full fixed-depth performance at 1.6B scale with a 3x smaller KV cache demonstrates that memory-efficient looped models are highly viable for consumer hardware and high-throughput production environments.

Who it affects

  • AI Developer
  • AI Researcher
  • Product Manager

How to use it

  1. 1Edge AI deployment: Run 1.6B-scale looped LLMs with 3x smaller KV cache footprint on memory-constrained devices.
  2. 2Low-latency chatbots: Speed up the prefill stage by up to 1.79x using a distilled student, reducing Time-to-First-Token.
  3. 3Efficient RLHF alignment: Compute gradients directly from saved rollout states during RL training to double the training throughput.

Limitations & caveats

  • The fixed-point convergence and prior learning mechanics increase model complexity, requiring careful hyperparameter tuning during training.
  • The evaluation is limited up to 1.6B parameters; scalability and stability on ultra-large scale models (e.g., 70B+) require further validation.

Related