Overcoming Test-Time Scaling Bottlenecks with Adaptive Looped Transformers: Introducing TaH2
自適應循環 Transformer 突破測試時縮放瓶頸:清華團隊提出 TaH2 架構
Looped Transformers reuse layers to save parameters but traditionally waste FLOPs by running fixed iterations on all tokens, causing them to underperform at matched compute. To address this, TaH2 jointly trains the model backbone with an iteration decider. Guided by lookahead depth supervision, TaH2 dynamically decides whether extra iterations will benefit a token's prediction. On challenging AIME benchmarks, TaH2 improves the accuracy-compute scaling slope by 53% over non-looped baselines and outperforms them by 3.4 points at matched compute.
Key points
Token-Adaptive Iteration
Instead of running fixed loop iterations for all tokens, TaH2 dynamically focuses extra computational resources on difficult tokens that actually benefit from looping.
Lookahead Depth Supervision
Jointly post-trains the backbone and an iteration decider using online lookahead labels to determine if further iterations yield better output quality.
Steeper Scaling Slope
On AIME benchmarks, TaH2 improves the accuracy-compute slope by 53% (2.74 vs 1.79) compared to the non-looped baseline.
Overcoming Performance Plateau
While existing looped models plateau with deeper iterations, TaH2's gain over baseline grows from +2.8 points at depth 2 to +3.9 points at depth 8.
How it works
| 非循環基準模型 (Non-Looped Baseline) | 現有循環模型 (Existing Looped) | TaH2 (自適應循環) | |
|---|---|---|---|
| Compute Allocation | 靜態 / Token 均等 | 靜態 / Token 均等(在簡單 token 上浪費算力) | 動態 / Token 自適應(依難易度分配算力) |
| AIME Scaling Slope | 1.79 | 斜率較陡,但在相同計算量下表現較差 | 2.74 (提升 53%) |
| Scaling with Deeper Iterations | 標準基準表現 | 趨於平緩、增長飽和 | 持續增長(深度 8 時領先 3.9 分) |
Why it matters
As LLMs increasingly rely on test-time scaling for complex reasoning, executing extra computation efficiently without wasting FLOPS on easy tokens is critical. TaH2 demonstrates that adaptive-depth looped transformers can outperform non-looped models at matched compute budgets. This provides a parameter- and compute-efficient pathway for designing next-generation reasoning models that scale during inference.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
How to use it
- 1Accelerating complex reasoning tasks (such as mathematics and code generation) by dynamically allocating compute budgets per token.
- 2Deploying parameter-efficient looped models on memory-constrained edge devices while preserving high-performance test-time scaling.
Limitations & caveats
- Requires joint post-training of the backbone and the iteration decider, which complicates the overall training pipeline.
- Varying token-level iteration depths can introduce engineering challenges for maintaining maximum parallel hardware utilization during batched inference.
Related

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude
艾倫人工智慧研究所開源 AstaBrief:比 Claude 快 3.5 倍的 8B 科學報告生成模型
Allen Institute for AI (Ai2) has open-sourced AstaBrief 8B, a specialized model for scientific report generation that achieves a 3.5x speedup over proprietary pipelines while maintaining high citation accuracy.
GALA: Distilling 3D Gaussian Avatars into Linear Blendshapes for Real-Time Animation
GALA:用線性混合變形蒸餾技術實現 3D Gaussian 虛擬化身即時動畫
GALA distills complex neural decoding of 3D Gaussian avatars into lightweight linear blendshapes, reducing CPU animation costs by up to 1000x and enabling 60fps real-time performance on mobile devices.
ScholarCatalyst: A Benchmark for Testing AI's Intuition in Retrieving Inspiring Research Papers
ScholarCatalyst:評估 AI 是否擁有「科學家直覺」的學術文獻檢索基準
ScholarCatalyst is a novel benchmark featuring annotations from 184 lead authors to evaluate whether AI can retrieve key inspiring papers from past literature based only on an initial research question.