Aivora
arXivVideo AIAdvanced

Breaking the Uniformity Trap: Scaling Video Diffusion Models via SplitMoE

突破均勻分佈陷阱:透過 SplitMoE 解決影片擴散模型的擴展瓶頸

2 min read
Breaking the Uniformity Trap: Scaling Video Diffusion Models via SplitMoE
The 30-second version

Applying traditional MoE to video diffusion often fails because forcing uniform token routing scatters spatially coherent video patches, leading to structural distortion. To address this "uniformity trap," SplitMoE introduces a split-role sparse architecture. It bifurcates the expert pool into "semantic experts" for high-level abstraction and "generic experts" for residual details, guided by prototype-based routing. Under equivalent computational budgets, SplitMoE achieves faster convergence, natural semantic clustering, and superior video generation quality.

Key points

01

The Uniformity Trap

Standard MoEs enforce uniform token routing, which conflicts with the spatiotemporal redundancy of video data, causing routing fragmentation and visual distortion.

02

Bifurcated Expert Roles

Explicitly divides the expert pool into "semantic experts" for high-level abstraction and "generic experts" for capturing residual visual details.

03

Prototype-Guided Routing

Replaces rigid balancing with prototype-guided routing and pull-push regularization, allowing tokens to cluster naturally by semantic similarity.

04

Coarse-to-Fine Denoising

SplitMoE uncovers an emergent coarse-to-fine denoising pattern, offering a promising, modality-aware scaling path for video world models.

How it works

Comparison of Traditional Visual MoE and SplitMoE
傳統視覺 MoE (Traditional MoE)SplitMoE (本研究提出)
Expert Pool同質專家池 (Homogeneous pool)異質雙重專家 (Semantic + Generic)
Routing Principle強制均勻平衡分流 (Enforced uniformity)原型引導與拉推非均勻分流 (Prototype-guided)
Visual Coherence易發生路由碎裂與結構失真自然依語意聚集,保持視覺連貫
Denoising Trajectory無明顯階段分工 (Unstructured)湧現由粗到細的去噪邏輯 (Coarse-to-fine)

Why it matters

As video generation and world models scale up, optimizing video diffusion models is critical. SplitMoE breaks the trend of blindly copying LLM-based MoE architectures, proving that a non-uniform, role-split sparse architecture is inherently better suited for multidimensional visual data. It reduces training and inference overhead while providing a clear blueprint for scaling large-scale video world models.

Who it affects

  • AI Developer
  • AI Researcher

How to use it

  1. 1Architectural design for high-quality, high-resolution video generation models
  2. 2Scaling computational efficiency for large-scale video world models

Limitations & caveats

  • The optimal ratio and allocation between semantic and generic experts may require hyperparameter tuning for different video datasets.
  • As a novel dynamic routing architecture, distributed training parallelization on standard hardware might require specialized engineering optimization.

Related

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
arXivVideo AI

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

超越時間線:利用「實體自傳」增強 AI 的長影片記憶與物體追蹤能力

This paper introduces Grounded Entity Biographies (GEB), a framework that compiles cross-clip visual observations of physical objects into retrievable biographies, significantly enhancing long-video question answering.

2 min read