Telescopic Language Models: One Training Run for Endless Compute Budgets
伸縮自如的語言模型:單次訓練即可適應多種運算資源預算的 Telescopic LM
Deploying models for varying compute budgets traditionally requires separate training or compression. While Matryoshka Language Model Suites (MLMS) offer fixed exits, exiting at non-designated layers causes performance to fail. Telescopic Language Models (TLM) solve this by training a nested-capacity Transformer using stochastic prefix supervision alongside a full-capacity anchor. Tested on a 200M proxy scale using 20B FineWeb-Edu tokens, a single TLM run functions as a valid language model at all 20 of its layer depths, reducing the quality-budget area under the curve by 43-44% while matching full-capacity quality and requiring ~12% lower GPU training cost.
Key points
Continuous Capacity Continuum
A single training run produces a model that is fully functional at any layer prefix, offering a continuous range of compute-budget trade-offs.
Stochastic Prefix Supervision
Trains a randomly truncated layer prefix against next-token targets in tandem with a full-capacity anchor pass, making every intermediate depth mathematically viable.
Eliminating Non-Exit Failures
Unlike fixed-exit baselines where non-exit layers collapse to high perplexity (10^2-10^5), TLM maintains smooth, low perplexity across all depths.
Lower Training Overhead
Requires no architectural changes and runs with ~12% lower GPU training cost per run compared to fixed-exit suites.
How it works
| Matryoshka Suites (MLMS) | Telescopic LM (TLM) | |
|---|---|---|
| Supported Exit Depths | 僅限預先設定的少數固定出口層 | 任何層前綴(整個容量軸連續有效) |
| Non-Designated Layer Perplexity | 極差(困惑度達 10^2 - 10^5) | 維持高水準,效能隨深度平滑過渡 |
| GPU Training Cost | 基準 (100%) | 比 MLMS 降低約 12% |
| Quality-Budget AUC | 基準 | 面積減少約 43% - 44% |
Why it matters
This work demonstrates that the training objective, rather than architecture, is what makes a model truly elastic. It shifts the paradigm of variable-resource deployment. Instead of training multiple models or setting rigid exit points for different target devices, teams can train a single TLM and dynamically scale the active layer count on-the-fly based on server load or device capabilities, optimizing overall GPU fleet efficiency.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Dynamic load balancing: reduce inference layers dynamically during high-traffic spikes to maintain low latency, and use full capacity during off-peak hours.
- 2Single-weight cross-device deployment: distribute one model to edge, mobile, and cloud devices, letting each run at its hardware-optimal layer depth without conversion.
Limitations & caveats
- The evaluations were primarily conducted at a 200M proxy model scale, meaning scaling behavior on multi-billion parameter models remains to be fully proven.
- Although cheaper than MLMS, TLM still requires dual forward-backward passes per training step, carrying a higher computational cost than training a single standard static model.
Related
Minimally Invasive Steering of LMs: Optimizing Rewards Without Quality Degradation
微創型語言模型導向技術:利用 MISVO 在不損害生成品質下優化輸出
This paper introduces MISVO, a minimally invasive steering method that uses local KL geometry to optimize LLM outputs for test-time rewards without parameter updates or quality degradation.

Accelerating MoE Training for Biological Foundation Models with NVIDIA Transformer Engine
NVIDIA Transformer Engine 加速生物基礎模型 MoE 訓練:吞吐量提升達 2.21 倍
This guide demonstrates how to use NVIDIA BioNeMo and Transformer Engine's optimized primitives to overcome MoE training bottlenecks in biological models, boosting throughput by up to 2.21x.

LFM2.5-VL-DSpark: Accelerating Vision-Language Models with Minimal Overhead
LFM2.5-VL-DSpark:以超低開銷將多模態模型推論速度提升達 3 倍
Liquid AI has released a 280M parameter DSpark draft model for LFM2.5-VL-3B, boosting decoding speeds up to 3.13 on-device with only 8.9% parameter overhead.