Embedding Prediction Helps Image Generation: NEPA-DiT Achieves Superior FID with Much Less Compute
預測嵌入向量助力影像生成:全新 NEPA-DiT 框架以超低算力達到更優 FID
Traditional Diffusion Transformers (DiTs) embed class labels or text prompts once and reuse them statically throughout all denoising steps. This paper introduces Next-Embedding Predictive Autoregression (NEPA), which trains a Transformer to predict continuous embeddings. Through Multi-Embedding Prediction, the system dynamically predicts the embeddings of the clean image at each denoising step, adapting the conditioning signal to the current noisy state. The resulting NEPA-DiT-XL model achieves an outstanding FID of 1.32 on ImageNet 256x256 while using only a third of the training compute required by REPA.
Key points
Dynamic Conditioning
Unlike traditional DiTs that reuse static conditions, NEPA dynamically updates predicted embeddings at each denoising step to adapt to the current noise level.
Next-Embedding Prediction
Treats the clean image's embeddings as the sequence following the noisy input, training a Transformer to predict continuous embeddings.
Drastic Compute Reduction
Combined with REPA, NEPA-DiT-XL achieves a state-of-the-art FID of 1.32 while slashing training compute requirements to just about one-third.
How it works
Why it matters
This research opens a highly efficient pathway for diffusion model architectures. Training generative models is notoriously expensive; NEPA proves that dynamic guidance via predicted embeddings not only elevates image quality to a superb 1.32 FID but also slashes compute costs to one-third. This democratizes high-performance image generation for teams with limited computational budgets.
Who it affects
- AI Researcher
- AI Developer
- Enterprise Leader
How to use it
- 1High-efficiency conditional image generation
- 2Denoising process optimization in Diffusion Transformers (DiTs)
Limitations & caveats
- Increased inference latency, as a second prediction network (NEPA) must be run at every denoising step.
- Highly dependent on the semantic quality and continuity of the underlying pre-trained embedding space.
Related
MatLoom: Layered Text-to-Material Generation in a Compact Program Space
MatLoom:用極簡程式空間實現分層式文字生成 PBR 材質
MatLoom is a compact, layer-oriented language that enables LLMs to generate highly editable PBR materials as structured code, outperforming traditional diffusion models in prompt alignment.
Looped-DiT: Scaling Image Generation Efficiency via Recurrent Transformer Blocks
Looped-DiT:藉由循環 Transformer 區塊實現影像生成的高效運算擴展
Looped-DiT reuses shared Transformer blocks within each denoising step, enabling a 260M-parameter model to outperform a 6.5x larger counterpart with 4.9x less inference compute.
MGFlow: Unifying Distributional Training for High-Quality One-Step Visual Generation
統一分佈訓練框架 MGFlow:實現高品質單步視覺生成
This paper presents a unified theoretical framework for distributional training and introduces MGFlow, which uses Gaussian mixtures and Wasserstein gradient flows to achieve state-of-the-art one-step visual generation.