Aivora
arXivVideo AIAdvanced

PDMD: Stabilizing Video Diffusion Distillation via Projected Distribution Matching

PDMD:投影分佈匹配蒸餾技術,解決影片擴散模型的一步生成偽影

2 min read
PDMD: Stabilizing Video Diffusion Distillation via Projected Distribution Matching
The 30-second version

Modern video diffusion models are slow due to many denoising steps. While Distribution Matching Distillation (DMD) reduces steps (NFE) to just 4, it suffers from artifacts and oversaturation caused by accumulated critic errors. PDMD (Projected Distribution Matching Distillation) solves this by projecting DMD updates to filter out critic errors. Supported by mathematical proof in high dimensions, PDMD requires only a one-line code change with no extra networks or data. It delivers state-of-the-art visual and audio quality on Wan2.1 and MiniMax-H3 benchmarks.

Key points

01

Eliminating Critic Error

Identifies that video degradation and oversaturation in DMD stem from accumulated critic errors over student updates.

02

Orthogonal Projection

Projects DMD updates orthogonal to the student-critic endpoint residual, removing critic errors while preserving the training signal.

03

One-Line Implementation

Extremely simple to deploy, requiring only a single-line change to DMD without extra losses, networks, or data.

04

SOTA 4-NFE Benchmarks

Achieves 83.73 VBench score with Wan2.1 and 83.17 VideoGen-Eval score with MiniMax-H3, delivering superior video and audio quality.

How it works

PDMD Projected Filtering Flow
Noisy QueryStudent ModelCritic ModelEndpoint ResidualDMD Update VectorPDMD ProjectionStable Update

Why it matters

Distilling video models to ultra-low NFEs is crucial for real-time applications. PDMD provides an elegant, mathematically proven geometric filtering method that resolves training instability in rapid distillation. Since it introduces zero computational overhead, it is highly practical for large-scale production pipelines.

Who it affects

  • AI Developer
  • AI Researcher
  • Product Manager

How to use it

  1. 1Accelerated video generation distillation (e.g., 4-step Wan2.1)
  2. 2Joint video-audio synced generation (e.g., MiniMax-H3 distillation)

Limitations & caveats

  • It remains a distillation approach, requiring offline training on specific base models rather than zero-shot acceleration.
  • The mathematical proofs rely on high-dimensional assumptions, which might have less pronounced benefits in low-dimensional spaces.