Aivora
arXivImage AIAdvanced

MGFlow: Unifying Distributional Training for High-Quality One-Step Visual Generation

統一分佈訓練框架 MGFlow:實現高品質單步視覺生成

2 min read
MGFlow: Unifying Distributional Training for High-Quality One-Step Visual Generation
The 30-second version

One-step visual generation significantly speeds up inference but lacks unified theoretical guidance for distributional training. This study introduces a unifying framework that connects global objectives to pointwise updates via Wasserstein gradient flow. Leveraging this, the authors propose MGFlow, which models feature distributions using Gaussian mixtures with adjustable granularity. MGFlow addresses mode collapse using mass-constrained sample assignment and paired component updates. It achieves state-of-the-art results on ImageNet and successfully post-trains FLUX.2 [klein] 4B into a single-step generator that outperforms its original 4-step counterpart.

Key points

01

Unified Training Framework

Separates distribution modeling from matching discrepancy, bridging global objectives and pointwise updates via Wasserstein gradient flow.

02

Mixture Gaussian Flow

Models feature distributions using Gaussian mixtures, supporting optimal transport and score-based matching with adjustable granularity.

03

Resolving Mode Collapse

Combines mass-constrained sample assignment with paired component updates to effectively tackle mode collapse issues.

04

One-Step Outperforming Multi-Step

Achieves SOTA on ImageNet and post-trains FLUX.2 [klein] 4B into a 1-step model that beats the original 4-step version.

How it works

MGFlow Workflow and Optimization Architecture
Feature AlignmentFeature AlignmentAssign SamplesCompute Pointwise UpdateUpdate Generator ParamsReal FeaturesGenerated FeaturesGaussian MixtureMass-ConstrainedAssignmentWasserstein GradientFlowOptimized 1-StepGenerator

Why it matters

This research provides a solid mathematical foundation for accelerating text-to-image models through one-step generation. By overcoming the single-Gaussian limitations of traditional FD-Loss, MGFlow's flexible mixture approach yields impressive practical results. Successfully converting the 4B-parameter FLUX.2 into a 1-step generator that beats its 4-step counterpart marks a major milestone in practical, ultra-fast visual generation.

Who it affects

  • AI Developer
  • AI Researcher

How to use it

  1. 1Post-training large diffusion models (like FLUX) into ultra-fast, single-step generation models.
  2. 2Accelerating generation pipelines and reducing computational latency in high-resolution image synthesis tasks.

Limitations & caveats

  • Although MGFlow offers adjustable granularity, modeling and assigning Gaussian mixtures in extremely high-dimensional feature spaces poses computational complexity challenges.
  • The method relies on matching features within frozen representation spaces, meaning generation quality is still bounded by the representation capacity of the feature extractor.

Related