AI Daily ·
Unlocking Base Model Reasoning and Advancements in Multimodal Diffusion and Image Processing
今日 AI 重點
Today's AI research highlights significant breakthroughs across multiple domains. A new study reveals that specific starting tokens can unlock latent reasoning capabilities in base models. Meanwhile, an innovative framework allows researchers to monitor the internal states of multimodal diffusion transformers in real time. Additionally, an agentic pipeline has been introduced to automatically convert raster flowcharts into editable draw.io files, collectively improving model performance, interpretability, and practical usability.
- 01arXivLLM
Base Models Can Reason: Unlocking Latent Performance with Strategic Starting Tokens
This study demonstrates that base language models possess latent reasoning capabilities that can be unlocked by forcing specific starting token cues (such as ".\n\nOkay" or "Alright,"). For instance, Olmo-3-7B's MATH-500 accuracy jumps from 42% to 78% when cued. The authors show that RL primarily increases the likelihood of generating these existing cues. Through causal data interventions, they successfully turned arbitrary words like 'chicken' into reasoning triggers, proving that these cues map to specific document types in the pre-training data.
- 02arXivAI Research
Learning to Read Contextual Tokens in Diffusion Transformers
Multimodal Diffusion Transformers (MM-DiTs) evolve text and image tokens together, creating dynamic contextual tokens. This paper trains a lightweight bottleneck network to map these tokens into a frozen LLM's input space, allowing the LLM to answer questions about the emerging image. The study reveals that contextual tokens capture rich, global semantics surprisingly early in denoising, even under empty prompts. Leveraging these insights, the authors propose 'Contextual Alignment' to reinforce this semantic encoding, directly improving generation quality and coverage.
- 03arXivAI Agent
One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline
Repurposing flowchart figures for different formats (slides, papers, mobile) is tedious, and existing generative or parsing approaches often distort shapes or break graph connections. This paper proposes a multi-stage agentic pipeline (Parse, Style, Layout) where each stage pairs a generator agent with a critic combining visual-language feedback and deterministic checks. The output is a fully editable draw.io XML. Tested on a 100-flowchart benchmark across five aspect ratios, this system achieves 68.6% Content Fidelity, significantly outperforming prior baselines.
- 04arXivAI Research
BiasFlow: Geometric Monitoring and Backbone Regularization for Spurious Feature Reliance
While Worst-Group Accuracy (WGA) evaluates trained predictors, it fails to capture how a frozen backbone behaves under new heads. The authors present BiasFlow, a hook-based toolkit using geometric diagnostics (IBMI, W-IBMI) to monitor class-attribute alignment. Paired with BiasFlow Regularization (BFR), a class-conditional alignment penalty, it mitigates spurious feature reliance. BFR improves UrbanCars WGA by up to +26.0 pp and boosts frozen CelebA-Std backbone WGA from 40.7% to 64.1% when evaluated with fresh classification heads.
- 05arXivAI Research
Direct Intermediate Initialization for Tilted Diffusion Samplers
Sequential Monte Carlo (SMC) diffusion samplers like MCGDiff struggle under finite particle budgets, often failing to capture rare posterior modes. This paper proposes Direct Intermediate Initialization (DII). By pulling back intermediate Gaussian-tilted targets to a softened clean-space posterior, samples can be generated via an approximate solver (like MMPS) at an intermediate timestep, then analytically mapped to the noisy target using a Gaussian bridge. Running only the remaining SMC suffix trades asymptotic consistency for performance, yielding a 2x improvement in sliced Wasserstein distance on Gaussian mixtures and over a 10x improvement on rare-mode inverse problems.