Aivora
arXivLLMIntermediate

Telescopic Language Models: One Training Run for Endless Compute Budgets

伸縮自如的語言模型:單次訓練即可適應多種運算資源預算的 Telescopic LM

2 min read
Telescopic Language Models: One Training Run for Endless Compute Budgets
The 30-second version

Deploying models for varying compute budgets traditionally requires separate training or compression. While Matryoshka Language Model Suites (MLMS) offer fixed exits, exiting at non-designated layers causes performance to fail. Telescopic Language Models (TLM) solve this by training a nested-capacity Transformer using stochastic prefix supervision alongside a full-capacity anchor. Tested on a 200M proxy scale using 20B FineWeb-Edu tokens, a single TLM run functions as a valid language model at all 20 of its layer depths, reducing the quality-budget area under the curve by 43-44% while matching full-capacity quality and requiring ~12% lower GPU training cost.

Key points

01

Continuous Capacity Continuum

A single training run produces a model that is fully functional at any layer prefix, offering a continuous range of compute-budget trade-offs.

02

Stochastic Prefix Supervision

Trains a randomly truncated layer prefix against next-token targets in tandem with a full-capacity anchor pass, making every intermediate depth mathematically viable.

03

Eliminating Non-Exit Failures

Unlike fixed-exit baselines where non-exit layers collapse to high perplexity (10^2-10^5), TLM maintains smooth, low perplexity across all depths.

04

Lower Training Overhead

Requires no architectural changes and runs with ~12% lower GPU training cost per run compared to fixed-exit suites.

How it works

Comparison: Telescopic LM (TLM) vs. Fixed-Exit Suites (MLMS)
Matryoshka Suites (MLMS)Telescopic LM (TLM)
Supported Exit Depths僅限預先設定的少數固定出口層任何層前綴(整個容量軸連續有效)
Non-Designated Layer Perplexity極差(困惑度達 10^2 - 10^5)維持高水準,效能隨深度平滑過渡
GPU Training Cost基準 (100%)比 MLMS 降低約 12%
Quality-Budget AUC基準面積減少約 43% - 44%

Why it matters

This work demonstrates that the training objective, rather than architecture, is what makes a model truly elastic. It shifts the paradigm of variable-resource deployment. Instead of training multiple models or setting rigid exit points for different target devices, teams can train a single TLM and dynamically scale the active layer count on-the-fly based on server load or device capabilities, optimizing overall GPU fleet efficiency.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Dynamic load balancing: reduce inference layers dynamically during high-traffic spikes to maintain low latency, and use full capacity during off-peak hours.
  2. 2Single-weight cross-device deployment: distribute one model to edge, mobile, and cloud devices, letting each run at its hardware-optimal layer depth without conversion.

Limitations & caveats

  • The evaluations were primarily conducted at a 200M proxy model scale, meaning scaling behavior on multi-billion parameter models remains to be fully proven.
  • Although cheaper than MLMS, TLM still requires dual forward-backward passes per training step, carrying a higher computational cost than training a single standard static model.

Related