AI Daily ·
AI Agent Evaluation and Architecture Breakthroughs Drive Training Efficiency
今日 AI 重點
Today's AI highlights focus on agent evaluation and architectural efficiency. Hugging Face introduced ThinkingBox for backend-based agent evaluation and Holo4, a generalist computer-use agent for complex workflows. Additionally, KV-streams significantly accelerates agentic reinforcement learning training by streaming the KV cache, while TaH2 and TACT improve test-time scaling and communication efficiency, respectively.
01Hugging FaceAI AgentThe Agent Said It Was Done, the Database Disagreed: Introducing ThinkingBox
ThinkingBox is a stateful sandbox benchmark consisting of 507 business workflows. It reveals that in 67.24% of failures, agents terminated cleanly with valid tool calls but actually corrupted the backend database with wrong fields or extra side effects. By repeating each task 20 times, ThinkingBox measures both broad capability (pass@20) and absolute consistency (Observed 20/20), proving that models like Kimi-K3 have high breadth but low consistency compared to Claude Opus 5.5.
- 02arXivAI Research
KV-streams: Boosting Context Compaction Efficiency in Agentic Reinforcement Learning
Training agentic LLMs with long contexts requires context compaction to manage GPU memory. However, traditional compaction methods frequently flush and rebuild the KV cache (prefilling), severely slowing down training. To solve this, researchers introduced KV-streams. Instead of flushing the KV cache after compaction, it streams the cache forward. Compatible with multiple compaction strategies, KV-streams delivers a 2.6x to 5x training speedup. Intriguingly, the streamed KV cache functions as a recurrent state, carrying forward long-term memories lost from the active context—a capability that emerges naturally through RL alone.
- 03arXivAI Research
Overcoming Test-Time Scaling Bottlenecks with Adaptive Looped Transformers: Introducing TaH2
Looped Transformers reuse layers to save parameters but traditionally waste FLOPs by running fixed iterations on all tokens, causing them to underperform at matched compute. To address this, TaH2 jointly trains the model backbone with an iteration decider. Guided by lookahead depth supervision, TaH2 dynamically decides whether extra iterations will benefit a token's prediction. On challenging AIME benchmarks, TaH2 improves the accuracy-compute scaling slope by 53% over non-looped baselines and outperforms them by 3.4 points at matched compute.
- 04arXivAI Agent
Building Communication-Efficient Social Intelligence in Language Agents with TACT
Socially intelligent agents often struggle to coordinate concisely. The TACT (Teacher-Assisted Communication Training) framework addresses this by revising student-generated actions through two specialized teachers: an expression specialist (which reduces wordiness) and a strategy specialist (which optimizes tactical moves). By simulating partner reactions, TACT selects teacher references that balance local goal support against action-token costs, distilling this efficient social intelligence back into the student model for independent deployment.
- 05arXivImage AI
Copy the Same, Distill the Difference: A Smart Initialization Paradigm for Linear Vision Transformers
Linear Vision Transformers (Linear ViTs) offer linear efficiency but typically suffer from poor performance due to a lack of pre-trained initializations. This study reveals that direct attention copying fails because Softmax and linear attention weights are highly operator-specific. Conversely, MLP weights are operator-agnostic and carry core representation capabilities, meaning they can be directly copied. Combining direct MLP copying with routing-based attention distillation (CSDD) allows Linear ViTs to match or even outperform traditional Softmax ViTs.
06Hugging FaceAI AgentHolo4: The Generalist Computer-Use Agent Bridging GUI, Code, and APIs
H company has launched Holo4 (available in 27B dense and 35B MoE) and Holotron4 Nano. Unlike traditional single-interface agents, Holo4 natively combines GUI navigation, code execution, MCP, and API calls across desktops, web, Android, and sandboxes. It achieves highly competitive performance on OSWorld 2.0 and AutomationBench at a fraction of the cost of frontier closed-source models.
- 07arXivRobotics
Learning Contractive Dynamical Representations for Composite Adaptive Control
This paper presents a representation-learning framework for composite adaptive tracking control under coupled disturbances. By linking classical disturbance-accommodating control (DAC) with adaptive disturbance-rejection, the authors introduce a hard-EM procedure with a Kalman smoother to identify contractive dynamical representations of latent disturbances. Validated on a slippery ground vehicle with sloshing liquid and a pendulum, the method achieves superior tracking and predictive performance over LTI and PD baselines.