Aivora
arXivImage AIAdvanced

Embedding Prediction Helps Image Generation: NEPA-DiT Achieves Superior FID with Much Less Compute

預測嵌入向量助力影像生成:全新 NEPA-DiT 框架以超低算力達到更優 FID

2 min read
Embedding Prediction Helps Image Generation: NEPA-DiT Achieves Superior FID with Much Less Compute
The 30-second version

Traditional Diffusion Transformers (DiTs) embed class labels or text prompts once and reuse them statically throughout all denoising steps. This paper introduces Next-Embedding Predictive Autoregression (NEPA), which trains a Transformer to predict continuous embeddings. Through Multi-Embedding Prediction, the system dynamically predicts the embeddings of the clean image at each denoising step, adapting the conditioning signal to the current noisy state. The resulting NEPA-DiT-XL model achieves an outstanding FID of 1.32 on ImageNet 256x256 while using only a third of the training compute required by REPA.

Key points

01

Dynamic Conditioning

Unlike traditional DiTs that reuse static conditions, NEPA dynamically updates predicted embeddings at each denoising step to adapt to the current noise level.

02

Next-Embedding Prediction

Treats the clean image's embeddings as the sequence following the noisy input, training a Transformer to predict continuous embeddings.

03

Drastic Compute Reduction

Combined with REPA, NEPA-DiT-XL achieves a state-of-the-art FID of 1.32 while slashing training compute requirements to just about one-third.

How it works

NEPA Dynamic Conditioning Process Flow
InputCurrent statePredict clean embedDynamic conditionDenoise inputOutputLabel/PromptNoisy Image xtNEPA PredictorPredicted EmbedDiT GeneratorDenoised xt-1

Why it matters

This research opens a highly efficient pathway for diffusion model architectures. Training generative models is notoriously expensive; NEPA proves that dynamic guidance via predicted embeddings not only elevates image quality to a superb 1.32 FID but also slashes compute costs to one-third. This democratizes high-performance image generation for teams with limited computational budgets.

Who it affects

  • AI Researcher
  • AI Developer
  • Enterprise Leader

How to use it

  1. 1High-efficiency conditional image generation
  2. 2Denoising process optimization in Diffusion Transformers (DiTs)

Limitations & caveats

  • Increased inference latency, as a second prediction network (NEPA) must be run at every denoising step.
  • Highly dependent on the semantic quality and continuity of the underlying pre-trained embedding space.

Related

MatLoom: Layered Text-to-Material Generation in a Compact Program Space
arXivImage AI

MatLoom: Layered Text-to-Material Generation in a Compact Program Space

MatLoom:用極簡程式空間實現分層式文字生成 PBR 材質

MatLoom is a compact, layer-oriented language that enables LLMs to generate highly editable PBR materials as structured code, outperforming traditional diffusion models in prompt alignment.

2 min read