Aivora
arXivAI ResearchAdvanced

LIFT: Breaking Transformer Feed-Forward Bottlenecks with Latent Information Feedback

LIFT 架構:透過老師監督引入潛在資訊回饋,突破 Transformer 的單向限制

2 min read
LIFT: Breaking Transformer Feed-Forward Bottlenecks with Latent Information Feedback
The 30-second version

Standard Transformers are purely feed-forward, passing information across steps only via the decoded token. LIFT (Latent Information Feedback Transformer) overcomes this bottleneck. During pretraining, LIFT pairs each input token with a dense latent state derived from a teacher model's distribution. LIFT is trained to predict both the next token and next state in a fully parallel manner. At inference, it feeds back its own predicted states. Experiments (135M to 1B parameters) show LIFT outperforms standard Transformers on reasoning and procedural tasks under token-matched budgets.

Key points

01

Breaking Feed-Forward Limitations

Traditional Transformers cannot pass deep layer representations back to shallow layers. LIFT enables deep-to-shallow latent state feedback to prevent redundant computation.

02

Teacher-Supervised State Learning

Leverages precomputed token distributions from an off-the-shelf LM as dense states, turning recurrent learning into a parallelizable teacher-forced task.

03

Parallel Pretraining & Low Inference Overhead

Pretraining remains fully parallel as states are precomputed; inference introduces a minor computational overhead that decreases with larger model sizes.

04

Excellent State Tracking & Data Efficiency

In state-tracking tasks, a tiny LIFT outperforms same-size standard Transformers trained on 8x more data.

How it works

Standard Transformer vs. LIFT
標準 TransformerLIFT
Information Flow僅透過解碼出的 TokenToken 搭配潛在狀態回饋
Deep-to-Shallow Path無(僅前饋傳播)有,透過預測狀態回饋
Pretraining Parallelism完全平行化完全平行化(利用預先計算的狀態)
State-Tracking Efficiency基準水準卓越(同效能僅需 1/8 資料量)

Why it matters

This work successfully integrates recurrent state-passing into Transformers without losing parallelizability during pretraining. By utilizing teacher supervision for latent state propagation, LIFT enables smaller models to match or exceed the reasoning and state-tracking performance of much larger, data-heavy traditional architectures. This opens up new pathways for designing highly efficient, resource-constrained LLMs.

Who it affects

  • AI Developer
  • AI Researcher

How to use it

  1. 1Procedural tasks and complex algorithm execution
  2. 2Long-term state tracking and multi-step logical reasoning applications

Limitations & caveats

  • Introduces a minor computational overhead at inference time, though it decreases as model scale increases.
  • Relies on an off-the-shelf, high-quality pretrained teacher model to precompute the dense states during training.

Related

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
arXivAI Research

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Imagine3D-LLM:讓多模態大模型在回答前「腦補」出 3D 場景

Inspired by human spatial reasoning, Imagine3D-LLM teaches multimodal LLMs to reconstruct multi-view images into a compact 3D Gaussian Splatting representation before answering, significantly improving spatial reasoning.

2 min read
Why Standard Metrics Fail: A Spectral Theory of LLM Graph Reconstruction
arXivAI Research

Why Standard Metrics Fail: A Spectral Theory of LLM Graph Reconstruction

為什麼標準指標不夠用?大語言模型圖形重建的譜理論與失真邊界

This paper proves that the Wasserstein distance of Laplacian spectra in LLM graph reconstruction is strictly bracketed by edge counts, revealing why aggregate metrics fail to capture complex structural editing.

2 min read