One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
單一區塊實現多重深度:具備深度編程專家庫的循環 Vision Transformer
Standard Vision Transformers rely on stacking multiple unique layers, leading to heavy memory footprints. reViT addresses this by recurrently running a single Transformer block. To retain depth-specific feature capacity, reViT introduces Depth-Programmed Experts, constructing the FFN at each recurrent layer as a weight-space convex combination of a shared expert bank guided by a depth coordinate. Evaluated on ImageNet-1k and DINOv2 distillation, reViT-B/16 matches DeiT III accuracy with ~70% fewer stored parameters, while supporting elastic-depth inference and dynamic or static target deployment.
Key points
Single-Block Recurrence
Recurrently reuses a single Transformer block, dramatically cutting storage memory footprint.
Weight-Space Merging
Represents depth-specific FFNs via convex weight-space combinations, outperforming token-dispatch MoE alternatives.
70% Parameter Reduction
reViT-B/16 achieves DeiT III accuracy on ImageNet-1k with ~70% fewer stored parameters.
Elastic Depth & Static Deployment
A single checkpoint supports multiple depth configurations dynamically or can be materialized offline as a standard dense graph.
How it works
Why it matters
For edge devices and resource-constrained platforms, storage footprint is a primary deployment bottleneck for Vision Transformers. reViT proves that recurrent single-block execution with depth-programmed weight merging can cut parameters by 70% without sacrificing representation quality or inference FLOPs efficiency. Its elastic-depth capability further allows a single trained checkpoint to adapt gracefully across various runtime latency and resource budgets.
Who it affects
- AI Researcher
- AI Developer
- Student & Learner
- Enterprise Leader
How to use it
- 1Memory-constrained edge devices and mobile vision backbones
- 2Adaptive compute inference scaling depth based on runtime load
- 3Efficient distilled feature extractors for segmentation and depth estimation
Limitations & caveats
- Inference FLOPs remain comparable to standard ViTs at matching depth, primary saving is storage rather than compute time
- Offline static materialization into dense graphs eliminates runtime routing overhead but increases deployment storage size
Related
Re-Evaluating AI Time Horizons: A Statistical Assessment of the METR Benchmark
重新審視 METR 時間跨度指標:以統計模型量化 AI 的真實任務能力
Re-analyzing METR benchmark data via splines and item-response theory reveals that AI task difficulty is non-linear, making a jump from 3 to 30 minutes far easier than from 30 minutes to 5 hours.
Building Persistent 3D Object Memory: How Ledger Tracks Objects from Egocentric Videos
打造過目不忘的 3D 空間記憶:Ledger 如何透過第一人稱影片追蹤隱形物體
Researchers introduce Ledger, a framework that builds a persistent 3D object memory from egocentric videos, significantly improving spatial question-answering accuracy for embodied agents.
Decoupling Exploration from Optimization: How ExpDis Boosts LLM Reasoning and Solution Diversity
探索與優化解耦:全新強化學習框架 ExpDis 提升大語言模型的推理多元性
The ExpDis framework decouples exploration from optimization in RLVR. By training explorers with novelty bonuses and distilling filtered trajectories into a student model, it prevents model degradation while fostering diverse reasoning.