Aivora
arXivAI ResearchIntermediate

One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts

單一區塊實現多重深度:具備深度編程專家庫的循環 Vision Transformer

2 min read
One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
The 30-second version

Standard Vision Transformers rely on stacking multiple unique layers, leading to heavy memory footprints. reViT addresses this by recurrently running a single Transformer block. To retain depth-specific feature capacity, reViT introduces Depth-Programmed Experts, constructing the FFN at each recurrent layer as a weight-space convex combination of a shared expert bank guided by a depth coordinate. Evaluated on ImageNet-1k and DINOv2 distillation, reViT-B/16 matches DeiT III accuracy with ~70% fewer stored parameters, while supporting elastic-depth inference and dynamic or static target deployment.

Key points

01

Single-Block Recurrence

Recurrently reuses a single Transformer block, dramatically cutting storage memory footprint.

02

Weight-Space Merging

Represents depth-specific FFNs via convex weight-space combinations, outperforming token-dispatch MoE alternatives.

03

70% Parameter Reduction

reViT-B/16 achieves DeiT III accuracy on ImageNet-1k with ~70% fewer stored parameters.

04

Elastic Depth & Static Deployment

A single checkpoint supports multiple depth configurations dynamically or can be materialized offline as a standard dense graph.

How it works

reViT Recurrent Block & Depth-Programmed Experts Architecture
Trajectory ControlExpert WeightsGenerates Depth FFNRecurrent InputLoop to Next DepthFinal Iteration DoneNormalized Depth CoordtShared Expert BankWeight-Space MergingSingle Recurrent BlockInput Token FeaturesFinal Output Features

Why it matters

For edge devices and resource-constrained platforms, storage footprint is a primary deployment bottleneck for Vision Transformers. reViT proves that recurrent single-block execution with depth-programmed weight merging can cut parameters by 70% without sacrificing representation quality or inference FLOPs efficiency. Its elastic-depth capability further allows a single trained checkpoint to adapt gracefully across various runtime latency and resource budgets.

Who it affects

  • AI Researcher
  • AI Developer
  • Student & Learner
  • Enterprise Leader

How to use it

  1. 1Memory-constrained edge devices and mobile vision backbones
  2. 2Adaptive compute inference scaling depth based on runtime load
  3. 3Efficient distilled feature extractors for segmentation and depth estimation

Limitations & caveats

  • Inference FLOPs remain comparable to standard ViTs at matching depth, primary saving is storage rather than compute time
  • Offline static materialization into dense graphs eliminates runtime routing overhead but increases deployment storage size

Related

Re-Evaluating AI Time Horizons: A Statistical Assessment of the METR Benchmark
arXivAI Research

Re-Evaluating AI Time Horizons: A Statistical Assessment of the METR Benchmark

重新審視 METR 時間跨度指標:以統計模型量化 AI 的真實任務能力

Re-analyzing METR benchmark data via splines and item-response theory reveals that AI task difficulty is non-linear, making a jump from 3 to 30 minutes far easier than from 30 minutes to 5 hours.

2 min read