Aivora
arXivAI ResearchAdvanced

Overcoming Test-Time Scaling Bottlenecks with Adaptive Looped Transformers: Introducing TaH2

自適應循環 Transformer 突破測試時縮放瓶頸:清華團隊提出 TaH2 架構

2 min read
Overcoming Test-Time Scaling Bottlenecks with Adaptive Looped Transformers: Introducing TaH2
The 30-second version

Looped Transformers reuse layers to save parameters but traditionally waste FLOPs by running fixed iterations on all tokens, causing them to underperform at matched compute. To address this, TaH2 jointly trains the model backbone with an iteration decider. Guided by lookahead depth supervision, TaH2 dynamically decides whether extra iterations will benefit a token's prediction. On challenging AIME benchmarks, TaH2 improves the accuracy-compute scaling slope by 53% over non-looped baselines and outperforms them by 3.4 points at matched compute.

Key points

01

Token-Adaptive Iteration

Instead of running fixed loop iterations for all tokens, TaH2 dynamically focuses extra computational resources on difficult tokens that actually benefit from looping.

02

Lookahead Depth Supervision

Jointly post-trains the backbone and an iteration decider using online lookahead labels to determine if further iterations yield better output quality.

03

Steeper Scaling Slope

On AIME benchmarks, TaH2 improves the accuracy-compute slope by 53% (2.74 vs 1.79) compared to the non-looped baseline.

04

Overcoming Performance Plateau

While existing looped models plateau with deeper iterations, TaH2's gain over baseline grows from +2.8 points at depth 2 to +3.9 points at depth 8.

How it works

Comparison: TaH2 vs Traditional Architectures
非循環基準模型 (Non-Looped Baseline)現有循環模型 (Existing Looped)TaH2 (自適應循環)
Compute Allocation靜態 / Token 均等靜態 / Token 均等(在簡單 token 上浪費算力)動態 / Token 自適應(依難易度分配算力)
AIME Scaling Slope1.79斜率較陡,但在相同計算量下表現較差2.74 (提升 53%)
Scaling with Deeper Iterations標準基準表現趨於平緩、增長飽和持續增長(深度 8 時領先 3.9 分)

Why it matters

As LLMs increasingly rely on test-time scaling for complex reasoning, executing extra computation efficiently without wasting FLOPS on easy tokens is critical. TaH2 demonstrates that adaptive-depth looped transformers can outperform non-looped models at matched compute budgets. This provides a parameter- and compute-efficient pathway for designing next-generation reasoning models that scale during inference.

Who it affects

  • AI Developer
  • AI Researcher
  • Product Manager

How to use it

  1. 1Accelerating complex reasoning tasks (such as mathematics and code generation) by dynamically allocating compute budgets per token.
  2. 2Deploying parameter-efficient looped models on memory-constrained edge devices while preserving high-performance test-time scaling.

Limitations & caveats

  • Requires joint post-training of the backbone and the iteration decider, which complicates the overall training pipeline.
  • Varying token-level iteration depths can introduce engineering challenges for maintaining maximum parallel hardware utilization during batched inference.

Related

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude
Hugging FaceAI Research

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude

艾倫人工智慧研究所開源 AstaBrief:比 Claude 快 3.5 倍的 8B 科學報告生成模型

Allen Institute for AI (Ai2) has open-sourced AstaBrief 8B, a specialized model for scientific report generation that achieves a 3.5x speedup over proprietary pipelines while maintaining high citation accuracy.

2 min read