Aivora

AI Daily ·

Pushing Generation Limits: Multimodal Self-Correction, Telescopic Language Models, and Efficient Visual Synthesis

今日 AI 重點

Today's AI research highlights major breakthroughs in model efficiency and generative capabilities. Telescopic Language Models (TLM) introduce a single training run to serve diverse compute budgets, while UMM-Reflection enables native image generation repair through interleaved reinforcement learning. Additionally, MGFlow and PDMD significantly advance efficient visual and video generation, and NstAgent successfully scales long-form story generation up to 100,000 words with narrative state tracking.

  1. Telescopic Language Models: One Training Run for Endless Compute Budgets
    01arXivLLM

    Telescopic Language Models: One Training Run for Endless Compute Budgets

    Deploying models for varying compute budgets traditionally requires separate training or compression. While Matryoshka Language Model Suites (MLMS) offer fixed exits, exiting at non-designated layers causes performance to fail. Telescopic Language Models (TLM) solve this by training a nested-capacity Transformer using stochastic prefix supervision alongside a full-capacity anchor. Tested on a 200M proxy scale using 20B FineWeb-Edu tokens, a single TLM run functions as a valid language model at all 20 of its layer depths, reducing the quality-budget area under the curve by 43-44% while matching full-capacity quality and requiring ~12% lower GPU training cost.

  2. Self-Correcting Multimodal Models: UMM-Reflection Enables Native Image Generation Repair via Interleaved RL
    02arXivAI Research

    Self-Correcting Multimodal Models: UMM-Reflection Enables Native Image Generation Repair via Interleaved RL

    Unified multimodal models can theoretically repair their own image generations, but training this iterative loop is difficult. UMM-Reflection addresses this by applying RL to complete reflection trajectories inside a single unified model. It uses 'sibling trajectories' sharing an initial image to compare strategies, using a single trajectory-level advantage to update both reflection tokens and image generation. This eliminates the need for external verifiers at inference time. It improves GenEval by 12.05 points over SFT and transfers well to other unseen benchmarks like WISE and T2I-CompBench++.

  3. MGFlow: Unifying Distributional Training for High-Quality One-Step Visual Generation
    03arXivImage AI

    MGFlow: Unifying Distributional Training for High-Quality One-Step Visual Generation

    One-step visual generation significantly speeds up inference but lacks unified theoretical guidance for distributional training. This study introduces a unifying framework that connects global objectives to pointwise updates via Wasserstein gradient flow. Leveraging this, the authors propose MGFlow, which models feature distributions using Gaussian mixtures with adjustable granularity. MGFlow addresses mode collapse using mass-constrained sample assignment and paired component updates. It achieves state-of-the-art results on ImageNet and successfully post-trains FLUX.2 [klein] 4B into a single-step generator that outperforms its original 4-step counterpart.

  4. FurE: 10x Faster 3D Animal Fur Reconstruction Without Animal Datasets
    04arXivAI Research

    FurE: 10x Faster 3D Animal Fur Reconstruction Without Animal Datasets

    Reconstructing realistic and editable animal fur from multi-view images is difficult due to self-occlusion and a lack of dedicated animal fur datasets. FurE overcomes this by optimizing a root-conditioned latent field decoded via a PCA decoder learned from human hair data. It reconstructs the underlying defurred animal body using a surface-constrained Gaussian Frosting representation and part-based priors. This approach generalizes across synthetic and real-world sequences, achieving a 10x training speedup over state-of-the-art dense per-strand optimization methods.

  5. PDMD: Stabilizing Video Diffusion Distillation via Projected Distribution Matching
    05arXivVideo AI

    PDMD: Stabilizing Video Diffusion Distillation via Projected Distribution Matching

    Modern video diffusion models are slow due to many denoising steps. While Distribution Matching Distillation (DMD) reduces steps (NFE) to just 4, it suffers from artifacts and oversaturation caused by accumulated critic errors. PDMD (Projected Distribution Matching Distillation) solves this by projecting DMD updates to filter out critic errors. Supported by mathematical proof in high dimensions, PDMD requires only a one-line code change with no extra networks or data. It delivers state-of-the-art visual and audio quality on Wan2.1 and MiniMax-H3 benchmarks.

  6. TokenCast: Accurate Token Consumption Forecasting for LLM Agents
    06arXivAI Agent

    TokenCast: Accurate Token Consumption Forecasting for LLM Agents

    LLM agent token usage is highly variable, often shifting by an order of magnitude due to dynamic tool feedback and compounding context growth. TokenCast solves this by learning composable cost representations for execution segments, tracking both immediate token use and future context expansion. It dynamically refreshes forecasts as the agent runs. TokenCast achieves an average of 14.5% prediction error reduction and saves 21.3% of token overhead in offline budget-control simulations.

  7. How to Loop MoE: Foil Architecture Flattens Experts and Unties Attention
    07arXivAI Research

    How to Loop MoE: Foil Architecture Flattens Experts and Unties Attention

    Looped Transformers reuse layers to maximize parameter utilization, while MoE limits computational cost per token. Combining them (Looped MoE) is highly promising but challenging. The authors propose Foil, which optimizes looped MoE under fixed parameters and compute. Foil (1) "flattens" the experts by halving expert layers, doubling experts per layer, and doubling the passes; and (2) "unties" the attention by assigning separate attention parameters to each pass while sharing experts/routers. At 100B tokens, Foil outperforms unflattened baselines, achieving lower pretraining loss and more balanced routing.

  8. Scaling Long-Form Story Generation via Narrative State Tracking (NstAgent)
    08arXivAI Agent

    Scaling Long-Form Story Generation via Narrative State Tracking (NstAgent)

    While LLMs excel in creative writing, maintaining narrative coherence over long texts remains a major challenge, with most methods limited to under 10K words. This paper presents Narrative State Tracking Agent (NstAgent), a training-free framework that structures and tracks characters, history, and future requirements. By testing on extended benchmarks from 10K to 100K words, the researchers demonstrated that NstAgent prevents the usual quality and consistency degradation in long-form generation.

Past issues