AI Daily ·
Pushing Generation Limits: Multimodal Self-Correction, Telescopic Language Models, and Efficient Visual Synthesis
今日 AI 重點
Today's AI research highlights major breakthroughs in model efficiency and generative capabilities. Telescopic Language Models (TLM) introduce a single training run to serve diverse compute budgets, while UMM-Reflection enables native image generation repair through interleaved reinforcement learning. Additionally, MGFlow and PDMD significantly advance efficient visual and video generation, and NstAgent successfully scales long-form story generation up to 100,000 words with narrative state tracking.
- 01arXivLLM
Telescopic Language Models: One Training Run for Endless Compute Budgets
Deploying models for varying compute budgets traditionally requires separate training or compression. While Matryoshka Language Model Suites (MLMS) offer fixed exits, exiting at non-designated layers causes performance to fail. Telescopic Language Models (TLM) solve this by training a nested-capacity Transformer using stochastic prefix supervision alongside a full-capacity anchor. Tested on a 200M proxy scale using 20B FineWeb-Edu tokens, a single TLM run functions as a valid language model at all 20 of its layer depths, reducing the quality-budget area under the curve by 43-44% while matching full-capacity quality and requiring ~12% lower GPU training cost.
- 02arXivAI Research
Self-Correcting Multimodal Models: UMM-Reflection Enables Native Image Generation Repair via Interleaved RL
Unified multimodal models can theoretically repair their own image generations, but training this iterative loop is difficult. UMM-Reflection addresses this by applying RL to complete reflection trajectories inside a single unified model. It uses 'sibling trajectories' sharing an initial image to compare strategies, using a single trajectory-level advantage to update both reflection tokens and image generation. This eliminates the need for external verifiers at inference time. It improves GenEval by 12.05 points over SFT and transfers well to other unseen benchmarks like WISE and T2I-CompBench++.
- 03arXivImage AI
MGFlow: Unifying Distributional Training for High-Quality One-Step Visual Generation
One-step visual generation significantly speeds up inference but lacks unified theoretical guidance for distributional training. This study introduces a unifying framework that connects global objectives to pointwise updates via Wasserstein gradient flow. Leveraging this, the authors propose MGFlow, which models feature distributions using Gaussian mixtures with adjustable granularity. MGFlow addresses mode collapse using mass-constrained sample assignment and paired component updates. It achieves state-of-the-art results on ImageNet and successfully post-trains FLUX.2 [klein] 4B into a single-step generator that outperforms its original 4-step counterpart.
- 04arXivAI Research
FurE: 10x Faster 3D Animal Fur Reconstruction Without Animal Datasets
Reconstructing realistic and editable animal fur from multi-view images is difficult due to self-occlusion and a lack of dedicated animal fur datasets. FurE overcomes this by optimizing a root-conditioned latent field decoded via a PCA decoder learned from human hair data. It reconstructs the underlying defurred animal body using a surface-constrained Gaussian Frosting representation and part-based priors. This approach generalizes across synthetic and real-world sequences, achieving a 10x training speedup over state-of-the-art dense per-strand optimization methods.
- 05arXivVideo AI
PDMD: Stabilizing Video Diffusion Distillation via Projected Distribution Matching
Modern video diffusion models are slow due to many denoising steps. While Distribution Matching Distillation (DMD) reduces steps (NFE) to just 4, it suffers from artifacts and oversaturation caused by accumulated critic errors. PDMD (Projected Distribution Matching Distillation) solves this by projecting DMD updates to filter out critic errors. Supported by mathematical proof in high dimensions, PDMD requires only a one-line code change with no extra networks or data. It delivers state-of-the-art visual and audio quality on Wan2.1 and MiniMax-H3 benchmarks.
- 06arXivAI Agent
TokenCast: Accurate Token Consumption Forecasting for LLM Agents
LLM agent token usage is highly variable, often shifting by an order of magnitude due to dynamic tool feedback and compounding context growth. TokenCast solves this by learning composable cost representations for execution segments, tracking both immediate token use and future context expansion. It dynamically refreshes forecasts as the agent runs. TokenCast achieves an average of 14.5% prediction error reduction and saves 21.3% of token overhead in offline budget-control simulations.
- 07arXivAI Research
How to Loop MoE: Foil Architecture Flattens Experts and Unties Attention
Looped Transformers reuse layers to maximize parameter utilization, while MoE limits computational cost per token. Combining them (Looped MoE) is highly promising but challenging. The authors propose Foil, which optimizes looped MoE under fixed parameters and compute. Foil (1) "flattens" the experts by halving expert layers, doubling experts per layer, and doubling the passes; and (2) "unties" the attention by assigning separate attention parameters to each pass while sharing experts/routers. At 100B tokens, Foil outperforms unflattened baselines, achieving lower pretraining loss and more balanced routing.
- 08arXivAI Agent
Scaling Long-Form Story Generation via Narrative State Tracking (NstAgent)
While LLMs excel in creative writing, maintaining narrative coherence over long texts remains a major challenge, with most methods limited to under 10K words. This paper presents Narrative State Tracking Agent (NstAgent), a training-free framework that structures and tracks characters, history, and future requirements. By testing on extended benchmarks from 10K to 100K words, the researchers demonstrated that NstAgent prevents the usual quality and consistency degradation in long-form generation.