Aivora
arXivImage AIIntermediate

Copy the Same, Distill the Difference: A Smart Initialization Paradigm for Linear Vision Transformers

複製相同、蒸餾相異:為線性視覺 Transformer 注入預訓練威力的全新初始化策略

2 min read
Copy the Same, Distill the Difference: A Smart Initialization Paradigm for Linear Vision Transformers
The 30-second version

Linear Vision Transformers (Linear ViTs) offer linear efficiency but typically suffer from poor performance due to a lack of pre-trained initializations. This study reveals that direct attention copying fails because Softmax and linear attention weights are highly operator-specific. Conversely, MLP weights are operator-agnostic and carry core representation capabilities, meaning they can be directly copied. Combining direct MLP copying with routing-based attention distillation (CSDD) allows Linear ViTs to match or even outperform traditional Softmax ViTs.

Key points

01

Direct MLP Copying

MLP weights are operator-agnostic; copying them directly from pre-trained Softmax ViTs preserves the majority of representation power.

02

Attention Copying Fails

Attention weights are operator-specific; directly copying them across Softmax-to-linear boundaries yields poor results, occasionally underperforming random initialization.

03

Distillation of Token Routing

A specialized distillation loss is used to guide the linear attention operator in recovering the token routing behavior of the Softmax version.

04

Broad Generalizability

The CSDD strategy consistently improves performance across various linear ViT variants, model sizes, and diverse datasets.

How it works

Copy the Same, Distill the Difference (CSDD) Architecture
SplitSplitDirect CopyRouting DistillAssembleAssemblePre-trained Softmax ViTMLP WeightsSoftmax AttentionDirect Copy MLPRouting DistilledAttentionEfficient Linear ViT

Why it matters

This work addresses the fundamental bottleneck of deploying efficient Linear ViTs by enabling weight reuse from foundation Softmax ViTs. By cleanly splitting the initialization into low-cost MLP copying and guided attention distillation, it removes the necessity for expensive from-scratch pre-training. This accelerates the deployment of high-resolution, low-latency vision models on resource-constrained devices.

Who it affects

  • AI Developer
  • AI Researcher
  • Product Manager

How to use it

  1. 1Edge and mobile deployment: Fast conversion of pre-trained Softmax ViTs to efficient Linear ViTs without losing accuracy.
  2. 2High-resolution image processing: Reducing memory footprints with linear attention while utilizing the initialized weights for fast training convergence.

Limitations & caveats

  • Unlike instant weight copying, this method still requires an active distillation training phase, which is not completely zero-cost.
  • The upper bound of the initialized model's performance remains constrained by the capacity of the teacher Softmax ViT.

Related

MatLoom: Layered Text-to-Material Generation in a Compact Program Space
arXivImage AI

MatLoom: Layered Text-to-Material Generation in a Compact Program Space

MatLoom:用極簡程式空間實現分層式文字生成 PBR 材質

MatLoom is a compact, layer-oriented language that enables LLMs to generate highly editable PBR materials as structured code, outperforming traditional diffusion models in prompt alignment.

2 min read