MGFlow: Unifying Distributional Training for High-Quality One-Step Visual Generation
統一分佈訓練框架 MGFlow:實現高品質單步視覺生成
One-step visual generation significantly speeds up inference but lacks unified theoretical guidance for distributional training. This study introduces a unifying framework that connects global objectives to pointwise updates via Wasserstein gradient flow. Leveraging this, the authors propose MGFlow, which models feature distributions using Gaussian mixtures with adjustable granularity. MGFlow addresses mode collapse using mass-constrained sample assignment and paired component updates. It achieves state-of-the-art results on ImageNet and successfully post-trains FLUX.2 [klein] 4B into a single-step generator that outperforms its original 4-step counterpart.
Key points
Unified Training Framework
Separates distribution modeling from matching discrepancy, bridging global objectives and pointwise updates via Wasserstein gradient flow.
Mixture Gaussian Flow
Models feature distributions using Gaussian mixtures, supporting optimal transport and score-based matching with adjustable granularity.
Resolving Mode Collapse
Combines mass-constrained sample assignment with paired component updates to effectively tackle mode collapse issues.
One-Step Outperforming Multi-Step
Achieves SOTA on ImageNet and post-trains FLUX.2 [klein] 4B into a 1-step model that beats the original 4-step version.
How it works
Why it matters
This research provides a solid mathematical foundation for accelerating text-to-image models through one-step generation. By overcoming the single-Gaussian limitations of traditional FD-Loss, MGFlow's flexible mixture approach yields impressive practical results. Successfully converting the 4B-parameter FLUX.2 into a 1-step generator that beats its 4-step counterpart marks a major milestone in practical, ultra-fast visual generation.
Who it affects
- AI Developer
- AI Researcher
How to use it
- 1Post-training large diffusion models (like FLUX) into ultra-fast, single-step generation models.
- 2Accelerating generation pipelines and reducing computational latency in high-resolution image synthesis tasks.
Limitations & caveats
- Although MGFlow offers adjustable granularity, modeling and assigning Gaussian mixtures in extremely high-dimensional feature spaces poses computational complexity challenges.
- The method relies on matching features within frozen representation spaces, meaning generation quality is still bounded by the representation capacity of the feature extractor.
Related

Meta Introduces Muse Image and Muse Video: Advancing Media Generation with Agentic Tools and Compute Scaling
Meta 發表 Muse Image 與 Muse Video:首款導入 Agentic 機制與推理期算力擴展的媒體生成模型
Meta launches Muse Image and previews Muse Video, pioneering agentic capabilities like tool use, self-refinement, and test-time compute scaling for media generation.