Copy the Same, Distill the Difference: A Smart Initialization Paradigm for Linear Vision Transformers
複製相同、蒸餾相異:為線性視覺 Transformer 注入預訓練威力的全新初始化策略
Linear Vision Transformers (Linear ViTs) offer linear efficiency but typically suffer from poor performance due to a lack of pre-trained initializations. This study reveals that direct attention copying fails because Softmax and linear attention weights are highly operator-specific. Conversely, MLP weights are operator-agnostic and carry core representation capabilities, meaning they can be directly copied. Combining direct MLP copying with routing-based attention distillation (CSDD) allows Linear ViTs to match or even outperform traditional Softmax ViTs.
Key points
Direct MLP Copying
MLP weights are operator-agnostic; copying them directly from pre-trained Softmax ViTs preserves the majority of representation power.
Attention Copying Fails
Attention weights are operator-specific; directly copying them across Softmax-to-linear boundaries yields poor results, occasionally underperforming random initialization.
Distillation of Token Routing
A specialized distillation loss is used to guide the linear attention operator in recovering the token routing behavior of the Softmax version.
Broad Generalizability
The CSDD strategy consistently improves performance across various linear ViT variants, model sizes, and diverse datasets.
How it works
Why it matters
This work addresses the fundamental bottleneck of deploying efficient Linear ViTs by enabling weight reuse from foundation Softmax ViTs. By cleanly splitting the initialization into low-cost MLP copying and guided attention distillation, it removes the necessity for expensive from-scratch pre-training. This accelerates the deployment of high-resolution, low-latency vision models on resource-constrained devices.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
How to use it
- 1Edge and mobile deployment: Fast conversion of pre-trained Softmax ViTs to efficient Linear ViTs without losing accuracy.
- 2High-resolution image processing: Reducing memory footprints with linear attention while utilizing the initialized weights for fast training convergence.
Limitations & caveats
- Unlike instant weight copying, this method still requires an active distillation training phase, which is not completely zero-cost.
- The upper bound of the initialized model's performance remains constrained by the capacity of the teacher Softmax ViT.
Related
Embedding Prediction Helps Image Generation: NEPA-DiT Achieves Superior FID with Much Less Compute
預測嵌入向量助力影像生成:全新 NEPA-DiT 框架以超低算力達到更優 FID
This study introduces NEPA, a framework that dynamically predicts and updates embedding conditions during denoising, enabling NEPA-DiT-XL to achieve an impressive 1.32 FID with only one-third of REPA's training compute.
MatLoom: Layered Text-to-Material Generation in a Compact Program Space
MatLoom:用極簡程式空間實現分層式文字生成 PBR 材質
MatLoom is a compact, layer-oriented language that enables LLMs to generate highly editable PBR materials as structured code, outperforming traditional diffusion models in prompt alignment.
Looped-DiT: Scaling Image Generation Efficiency via Recurrent Transformer Blocks
Looped-DiT:藉由循環 Transformer 區塊實現影像生成的高效運算擴展
Looped-DiT reuses shared Transformer blocks within each denoising step, enabling a 260M-parameter model to outperform a 6.5x larger counterpart with 4.9x less inference compute.