Breaking the Uniformity Trap: Scaling Video Diffusion Models via SplitMoE
突破均勻分佈陷阱:透過 SplitMoE 解決影片擴散模型的擴展瓶頸
Applying traditional MoE to video diffusion often fails because forcing uniform token routing scatters spatially coherent video patches, leading to structural distortion. To address this "uniformity trap," SplitMoE introduces a split-role sparse architecture. It bifurcates the expert pool into "semantic experts" for high-level abstraction and "generic experts" for residual details, guided by prototype-based routing. Under equivalent computational budgets, SplitMoE achieves faster convergence, natural semantic clustering, and superior video generation quality.
Key points
The Uniformity Trap
Standard MoEs enforce uniform token routing, which conflicts with the spatiotemporal redundancy of video data, causing routing fragmentation and visual distortion.
Bifurcated Expert Roles
Explicitly divides the expert pool into "semantic experts" for high-level abstraction and "generic experts" for capturing residual visual details.
Prototype-Guided Routing
Replaces rigid balancing with prototype-guided routing and pull-push regularization, allowing tokens to cluster naturally by semantic similarity.
Coarse-to-Fine Denoising
SplitMoE uncovers an emergent coarse-to-fine denoising pattern, offering a promising, modality-aware scaling path for video world models.
How it works
| 傳統視覺 MoE (Traditional MoE) | SplitMoE (本研究提出) | |
|---|---|---|
| Expert Pool | 同質專家池 (Homogeneous pool) | 異質雙重專家 (Semantic + Generic) |
| Routing Principle | 強制均勻平衡分流 (Enforced uniformity) | 原型引導與拉推非均勻分流 (Prototype-guided) |
| Visual Coherence | 易發生路由碎裂與結構失真 | 自然依語意聚集,保持視覺連貫 |
| Denoising Trajectory | 無明顯階段分工 (Unstructured) | 湧現由粗到細的去噪邏輯 (Coarse-to-fine) |
Why it matters
As video generation and world models scale up, optimizing video diffusion models is critical. SplitMoE breaks the trend of blindly copying LLM-based MoE architectures, proving that a non-uniform, role-split sparse architecture is inherently better suited for multidimensional visual data. It reduces training and inference overhead while providing a clear blueprint for scaling large-scale video world models.
Who it affects
- AI Developer
- AI Researcher
How to use it
- 1Architectural design for high-quality, high-resolution video generation models
- 2Scaling computational efficiency for large-scale video world models
Limitations & caveats
- The optimal ratio and allocation between semantic and generic experts may require hyperparameter tuning for different video datasets.
- As a novel dynamic routing architecture, distributed training parallelization on standard hardware might require specialized engineering optimization.
Related

NVIDIA VSS Blueprint 3.3: Lowering Visual AI Agent Costs with Smart Sampling and Agent Skills
NVIDIA VSS Blueprint 3.3 登場:以智慧採樣與 AI 代理技能,大幅降低視覺 AI 部署成本
NVIDIA VSS Blueprint 3.3 reduces development and runtime costs for visual AI agents through a prompt-based agent builder and Adaptive EVS token pruning.
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
超越時間線:利用「實體自傳」增強 AI 的長影片記憶與物體追蹤能力
This paper introduces Grounded Entity Biographies (GEB), a framework that compiles cross-clip visual observations of physical objects into retrievable biographies, significantly enhancing long-video question answering.
PDMD: Stabilizing Video Diffusion Distillation via Projected Distribution Matching
PDMD:投影分佈匹配蒸餾技術,解決影片擴散模型的一步生成偽影
PDMD stabilizes video diffusion distillation by projecting out critic errors using the student-critic endpoint residual, achieving high-quality 4-step generation with a one-line code change.