Aivora

AI Daily ·

Breakthroughs in Robotic World Models and Multimodal AI

今日 AI 重點

Today's AI landscape features major advancements, highlighted by RoboJEPA establishing scaling laws for robotic planning using an 8B-parameter latent world model. Meanwhile, Liquid AI unveiled open d1, an ultra-fast edge decision model supporting single-pass multimodal inference. Additionally, Microsoft introduced Agent Lightning v1.0 to optimize agentic reinforcement learning in real-world harnesses, driving new progress in embodied AI and intelligent agents.

  1. RoboJEPA: Scaling Robotic Latent World Models
    01arXivRobotics

    RoboJEPA: Scaling Robotic Latent World Models

    RoboJEPA is an 8B-parameter robotic world model trained on real-world data across 12 different robot embodiments. By applying the Joint Embedding Predictive Architecture (JEPA), the researchers established that the model's 'imagination error' follows a second-order power law relative to compute. This error acts as a reliable proxy for physical deployment performance, enabling predictable downstream planning and successful zero-shot long-horizon execution on real hardware.

  2. Decoupling Exploration from Optimization: How ExpDis Boosts LLM Reasoning and Solution Diversity
    02arXivAI Research

    Decoupling Exploration from Optimization: How ExpDis Boosts LLM Reasoning and Solution Diversity

    Standard Reinforcement Learning with Verifiable Rewards (RLVR) struggles to incentivize novel reasoning without degrading overall model quality. To address this, researchers developed the Exploration-Distillation (ExpDis) framework, which decouples exploration from optimization. ExpDis trains 'explorer' policies with novelty bonuses, filters their trajectories for correctness, and distills high-quality data into a 'student' policy. This iterative process outperforms DAPO across seven mathematical reasoning benchmarks and significantly improves pass@k scaling.

  3. EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory
    03arXivLLM

    EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory

    Traditional LLM editing often struggles with parametric side-effects. EngramEdit introduces a decoupled update method for conditional memory architectures (like DeepSeek Engram). It first identifies target memory representations for updated facts, then jointly updates shared n-gram embeddings. Crucially, it heavily penalizes updates to highly reused embeddings to prevent side-effects. Experiments show near-perfect edit success, a 3x accuracy improvement in CoT multi-hop reasoning over baselines, and excellent preservation of unrelated knowledge.

  4. Long-WAM: Scaling Context for Real-Time World-Action Models
    04arXivRobotics

    Long-WAM: Scaling Context for Real-Time World-Action Models

    This paper presents Long-WAM, a framework designed to scale visual history context for robotic world-action models under strict real-time constraints. The authors reveal that longer history only benefits control performance when the video foundation is pretrained autoregressively (AR). By first learning causal prediction from unlabeled videos and preserving this structure during action fine-tuning, combined with streaming encoding and asynchronous execution, Long-WAM achieves a low inference latency of 107.4 ms on RTX 5090. It enables highly dynamic manipulation, such as dynamic cup stacking, on physical platforms.

  5. "Rephrase Before You Act": Mitigating the Extreme Language Sensitivity of Vision-Language-Action Models
    05arXivRobotics

    "Rephrase Before You Act": Mitigating the Extreme Language Sensitivity of Vision-Language-Action Models

    Vision-Language-Action (VLA) models suffer from extreme sensitivity to phrasing—for instance, changing "switch on the stove" to "switch on the hot plate" crashes success rates from 100% to 2%. To address this, researchers introduced "Rephrase Before You Act." By evaluating various phrasings on a few training tasks, they leverage an LLM to distill 10-20 robust rewriting rules. At runtime, incoming instructions are systematically rephrased before hitting the frozen VLA, boosting relative success rates for $π_0$ by 16% to 27% without policy retraining.

  6. SciExam for ENSO: Evaluating AI Agents' Ability to Build Unsolved Climate Models
    06arXivAI Research

    SciExam for ENSO: Evaluating AI Agents' Ability to Build Unsolved Climate Models

    Evaluating scientific AI is difficult when there is no pre-existing ground truth. The SciExam for ENSO benchmark challenges AI agents to build low-order stochastic models of the El Niño-Southern Oscillation within a 6-hour window, relying only on self-written frozen diagnostics for feedback. Out of 12 agent systems tested, 6 successfully built models that outperformed an established, human-published model. Intriguingly, these agent-generated models naturally aligned with competing theories in an ongoing scientific debate regarding ENSO's asymmetry.

  7. RECAST: Active Evidence Construction via Adaptive Routing and Computation
    07arXivAI Agent

    RECAST: Active Evidence Construction via Adaptive Routing and Computation

    Traditional RAG struggles when answers require aggregating or computing data across heterogeneous sources. RECAST solves this by modeling evidence construction as a sequential decision process. A lightweight RouterLM iteratively selects primitive operations or tasks a frozen CompilerLM to generate executable code. Once sufficient evidence is derived and accepted, a frozen AnswerLM outputs the final answer. Trained via SFT and GRPO, RECAST outperforms strong baselines by up to 15.9% and demonstrates exceptional zero-shot generalization across held-out tasks.

  8. EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution
    08arXivRobotics

    EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution

    Robot foundation models often degrade when task instructions or environments change, while retraining via teleoperation is highly expensive. EmbodiedRSI overcomes this with a Fast-Slow Dual-System. It maintains competing code and skill hypotheses within a Hypothesis Graph, runs physical trials optimized by Value-of-Information (VoI), and performs Code-Skill Co-Evolution. Guided by Hierarchical Memory, it achieves 77% success on RoboCasa365 and zero-shot transfers to real robots with 71.3% success.

  9. Liquid AI Unveils open d1: Ultra-Fast Edge Decision Models with Single-Pass Multimodal Inference
    09Hugging FaceOpen Source

    Liquid AI Unveils open d1: Ultra-Fast Edge Decision Models with Single-Pass Multimodal Inference

    Liquid AI has introduced open d1, a family of edge-focused decision models built on Liquid Foundation Models (LFMs). Instead of generating tokens step-by-step, these models run in a single forward pass to output structured decisions directly, drastically reducing latency. d1-3B supports text and images, answering in under 50ms on edge devices like NVIDIA Jetson. d1-omni-600M is an early research release supporting text-image or text-audio modalities, beating Decider 2B with just a fraction of its parameters.

  10. Microsoft Releases Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Real-Harness Training
    10Microsoft ResearchAI Agent

    Microsoft Releases Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Real-Harness Training

    Traditional agentic RL requires rebuilding complex agent loops inside training frameworks, causing high development costs and behavioral drift. Agent Lightning v1.0 solves this by introducing 'Harnessed Agentic RL,' which inserts an LLM proxy between the agent and the model, allowing the deployment harness to participate directly in training. Built on just 3,500 lines of code, the framework natively supports standard Kubernetes jobs to avoid costly commercial sandboxes. It features 'Collocated Async RL' for GPU sharing, doubling end-to-end training speed. Using only 6,000 samples, it boosted Qwen3.5-9B's SWE-bench Verified Pass@1 score from 41.8% to 56.4%.

Past issues