LIFT: Breaking Transformer Feed-Forward Bottlenecks with Latent Information Feedback
LIFT 架構:透過老師監督引入潛在資訊回饋,突破 Transformer 的單向限制
Standard Transformers are purely feed-forward, passing information across steps only via the decoded token. LIFT (Latent Information Feedback Transformer) overcomes this bottleneck. During pretraining, LIFT pairs each input token with a dense latent state derived from a teacher model's distribution. LIFT is trained to predict both the next token and next state in a fully parallel manner. At inference, it feeds back its own predicted states. Experiments (135M to 1B parameters) show LIFT outperforms standard Transformers on reasoning and procedural tasks under token-matched budgets.
Key points
Breaking Feed-Forward Limitations
Traditional Transformers cannot pass deep layer representations back to shallow layers. LIFT enables deep-to-shallow latent state feedback to prevent redundant computation.
Teacher-Supervised State Learning
Leverages precomputed token distributions from an off-the-shelf LM as dense states, turning recurrent learning into a parallelizable teacher-forced task.
Parallel Pretraining & Low Inference Overhead
Pretraining remains fully parallel as states are precomputed; inference introduces a minor computational overhead that decreases with larger model sizes.
Excellent State Tracking & Data Efficiency
In state-tracking tasks, a tiny LIFT outperforms same-size standard Transformers trained on 8x more data.
How it works
| 標準 Transformer | LIFT | |
|---|---|---|
| Information Flow | 僅透過解碼出的 Token | Token 搭配潛在狀態回饋 |
| Deep-to-Shallow Path | 無(僅前饋傳播) | 有,透過預測狀態回饋 |
| Pretraining Parallelism | 完全平行化 | 完全平行化(利用預先計算的狀態) |
| State-Tracking Efficiency | 基準水準 | 卓越(同效能僅需 1/8 資料量) |
Why it matters
This work successfully integrates recurrent state-passing into Transformers without losing parallelizability during pretraining. By utilizing teacher supervision for latent state propagation, LIFT enables smaller models to match or exceed the reasoning and state-tracking performance of much larger, data-heavy traditional architectures. This opens up new pathways for designing highly efficient, resource-constrained LLMs.
Who it affects
- AI Developer
- AI Researcher
How to use it
- 1Procedural tasks and complex algorithm execution
- 2Long-term state tracking and multi-step logical reasoning applications
Limitations & caveats
- Introduces a minor computational overhead at inference time, though it decreases as model scale increases.
- Relies on an off-the-shelf, high-quality pretrained teacher model to precompute the dense states during training.
Related
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
Imagine3D-LLM:讓多模態大模型在回答前「腦補」出 3D 場景
Inspired by human spatial reasoning, Imagine3D-LLM teaches multimodal LLMs to reconstruct multi-view images into a compact 3D Gaussian Splatting representation before answering, significantly improving spatial reasoning.
The Convergence of Local Denoising Breakdown and Semantic Speciation in Generative Models
區域去噪失效與語意分化的同步:生成模型中的「相變」理論研究
This paper investigates why semantic class commitment and the breakdown of local denoising occur concurrently in generative models, proving that semantic information acts as their shared common cause.
Why Standard Metrics Fail: A Spectral Theory of LLM Graph Reconstruction
為什麼標準指標不夠用?大語言模型圖形重建的譜理論與失真邊界
This paper proves that the Wasserstein distance of Laplacian spectra in LLM graph reconstruction is strictly bracketed by edge counts, revealing why aggregate metrics fail to capture complex structural editing.