AI Daily ·
Technical Breakthroughs in Linear Attention and AI Agent Architectures
今日 AI 重點
Today's AI Daily highlights major advancements in model efficiency and agent architectures. In model optimization, STEPQuant and LeapQuant introduce innovative quantization techniques to significantly compress memory and accelerate inference for linear attention mechanisms. Concurrently, agentic frameworks are enhanced through meta-reasoning approaches for long-horizon tasks, alongside NVIDIA's release of Kumo Tabular for zero-shot tabular prediction, collectively pushing the boundaries of AI performance and computational efficiency.
- 01arXivLLM
STEPQuant: Spatial-Temporal Quantization for Linear Attention Recurrent States
Linear attention models replace growing KV caches with fixed-size recurrent states, which still pose memory bottlenecks during high-concurrency serving. Directly quantizing these states causes severe error accumulation over time. STEPQuant solves this by analyzing errors across two dimensions: temporally (allocating precision based on memory lifetime) and spatially (fitting scales based on key-row and value-column importance). It closely matches FP32 accuracy at 6-bit, achieves over 5x state compression, and reduces total serving memory by up to 68.7% when integrated into SGLang.
- 02arXivLLM
LeapQuant: Near-Lossless 8-Bit Recurrent State Quantization for Linear Attention LLMs
While hybrid LLMs utilize linear attention to compress context into a fixed recurrent state, the frequent read/write cycles of this state bottleneck inference. LeapQuant introduces a training-free 8-bit quantization framework to address this. It implements "per-window quantization" to jump over token windows and reduce error accumulation, keeps prominent outliers as high-precision "Compensator Tokens", and smooths the remaining residuals. Evaluated on Qwen, Kimi, and GLM models, LeapQuant delivers up to 1.47x end-to-end speedups on modern GPUs (including NVIDIA B200 and RTX 5090) with near-FP32 accuracy.
- 03arXivAI Research
LIFT: Breaking Transformer Feed-Forward Bottlenecks with Latent Information Feedback
Standard Transformers are purely feed-forward, passing information across steps only via the decoded token. LIFT (Latent Information Feedback Transformer) overcomes this bottleneck. During pretraining, LIFT pairs each input token with a dense latent state derived from a teacher model's distribution. LIFT is trained to predict both the next token and next state in a fully parallel manner. At inference, it feeds back its own predicted states. Experiments (135M to 1B parameters) show LIFT outperforms standard Transformers on reasoning and procedural tasks under token-matched budgets.
- 04arXivAI Agent
Thinking Before Thinking: Scaling AI Agents with Meta-Reasoning
As AI agents face longer and more complex tasks, controlling execution paths becomes critical. This paper proposes 'agentic meta-reasoning', which separates execution into task-level 'workers' and a high-level 'controller'. Instead of replaying full histories, the controller carries a compact summary, evaluates options against the remaining budget, and dispatches work using persistent memory. Across long-horizon benchmarks, this approach continued to scale and improve performance where direct control methods typically plateaued.
- 05arXivAI Agent
Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
While traditional approaches focus on upgrading an agent's reasoning, this paper explores environment design. Under frozen weights, a Builder model constructs an execution harness for a Target agent. By introducing Meta-Skills—rules on when to support and what resources to provide—the Builder learns from Target feedback. Evaluated on Harness-Bench and NewtonBench, this approach outperforms no-skill setups by 8.95 percentage points and direct-delivery baselines by 12.02 percentage points, showing a path to system-level self-improvement.
- 06arXivVideo AI
Breaking the Uniformity Trap: Scaling Video Diffusion Models via SplitMoE
Applying traditional MoE to video diffusion often fails because forcing uniform token routing scatters spatially coherent video patches, leading to structural distortion. To address this "uniformity trap," SplitMoE introduces a split-role sparse architecture. It bifurcates the expert pool into "semantic experts" for high-level abstraction and "generic experts" for residual details, guided by prototype-based routing. Under equivalent computational budgets, SplitMoE achieves faster convergence, natural semantic clustering, and superior video generation quality.
07Hugging FaceAI ResearchNVIDIA Introduces Kumo Tabular: An Open Foundation Model for Zero-Shot Tabular Prediction
NVIDIA released Kumo Tabular, an open foundation model for tabular data (ranging from 28M to 215M parameters). Applying the in-context learning paradigm of LLMs to structured data, it predicts labels for new rows in a single forward pass without task-specific training. Pretrained entirely on synthetic tables generated via Structural Causal Models (SCM), Kumo Tabular achieves state-of-the-art accuracy on TabArena, BeyondArena, and other benchmarks while running up to 17x faster than competitors.
08NVIDIA DeveloperOpen SourceDesigning AI-Native Software: Lessons from NVIDIA TensorRT Model Connect
NVIDIA TensorRT Model Connect is an open-source C++ project designed to make high-performance inference accessible. Instead of using AI agents as mere helper tools, the team built an "AI-native" architecture from scratch. By breaking tasks horizontally, defining strict outcomes rather than rigid prompts, isolating model families, and implementing strict validation, they successfully scaled the project to support 128 model families tested on NVIDIA GB300 as of July 2026.
- 09arXivRobotics
Skill-Space Shooting: Autonomous Robot Policy Improvement Guided by Foundation Models
Traditional robot policy correction relies heavily on human demonstrations. This paper proposes 'skill-space shooting,' which leverages foundation models to guide exploration over a set of reusable, short-horizon skills. When a robot encounters a failure, the system searches ('shoots') within this skill space to find a successful recovery path, turning successful trials into policy improvement data and enabling autonomous learning.
- 10arXivAI Research
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
While Multimodal Large Language Models (MLLMs) excel at single-image tasks, they struggle to integrate multi-view images for 3D spatial reasoning. To address this, researchers introduced Imagine3D-LLM. Inspired by how humans construct coarse 3D layouts mentally, this model appends learnable "summary tokens" after image tokens and decodes them into a compact 3D Gaussian Splatting (3DGS) representation. By training jointly with photometric reconstruction loss and standard next-token prediction, it propagates 3D-aware signals through its features, achieving superior performance on 3D understanding benchmarks.