Aivora

AI Daily ·

Breakthroughs in Agent Self-Improvement, Command Generation, and Efficient Fine-Tuning

今日 AI 重點

Today's AI highlights showcase major advancements in agent capabilities and efficient training. The RPG framework enables embodied agents to boost task success to 95% without weight updates through simulation practice and failure diagnosis. Meanwhile, KaliBench empowers an 8B model to rival massive MoE models in Kali Linux command generation. Additionally, the TACO optimizer drastically reduces memory requirements, enabling full-parameter fine-tuning of 32B models on a single GPU.

  1. KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux
    01arXivAI Agent

    KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux

    Deploying LLMs in security workflows requires precise translation of intent into strict CLI commands. KaliBench addresses this with 8,504 query-command pairs spanning 1,642 Kali Linux tools. Evaluations reveal that open-weight models struggle in this domain, with none exceeding 42% exact accuracy without explicit tool hints. To solve this, KaliBench introduces runtime-free verifiable rewards for training. Implementing this mechanism via SFT and reinforcement learning significantly boosts an 8B model to perform on par with a 685B MoE model.

  2. RPG Framework: Guided Self-Improvement Boosts Embodied Agent Success to 95% Without Weight Updates
    02arXivRobotics

    RPG Framework: Guided Self-Improvement Boosts Embodied Agent Success to 95% Without Weight Updates

    Developing reliable robot skills usually requires extensive human engineering. The RPG (Reconstruct, Practice, Go Real) framework automates this by extracting tasks from offline data, practicing in simulation, diagnosing failures via privileged states, and generating/refining symbolic skills and system prompts without updating neural network weights. RPG boosted task success on 22 manipulation tasks from 28.6% to 95.0%, outperforming GPT-6-based agents, and achieved a 100% success rate in real-world physical tests.

  3. VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness
    03arXivAI Agent

    VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness

    Developed by a team including Kaiming He, VISTA is a general-purpose visual harness that grants multimodal models long-horizon vision. It records raw observations in a lossless visual memory, allowing agents to actively retrieve and reorganize visual inputs during active reasoning. Tested on ARC-AGI-3, VISTA boosted Claude Opus 5.0's score to a perfect 100.00, completing games using 57.4% fewer actions than novice humans, proving its high generalization potential across visual puzzles.

  4. TACO Optimizer: Unleashing Full-Parameter 32B LLM Fine-Tuning on a Single GPU
    04arXivLLM

    TACO Optimizer: Unleashing Full-Parameter 32B LLM Fine-Tuning on a Single GPU

    Full-parameter LLM fine-tuning is heavily bottlenecked by the massive state memory of optimizers like AdamW. This paper introduces TACO, which computes updates by selecting only the sign of the absolute-maximum gradient entry in each column of the weight matrix. On OPT-13B, TACO reduces persistent optimizer states by 174x (from 27.7 GB to 0.16 GB) and peak training memory by 2.9x compared to 8-bit AdamW. Maintaining comparable accuracy and runtime, TACO enables full-parameter fine-tuning of 30-32B models on a single 80GB H100 GPU.

  5. Hierarchical Continuous Diffusion Language Models: Coupling Discrete Tokens with Continuous Latents
    05arXivLLM

    Hierarchical Continuous Diffusion Language Models: Coupling Discrete Tokens with Continuous Latents

    Traditional discrete diffusion language models suffer from independent token sampling during parallel decoding, while continuous alternatives lack structural constraints to guarantee valid token configurations. To resolve this, researchers introduced Hierarchical Continuous Diffusion Language Models (HC-DLM). HC-DLM couples discrete token generation with a continuous latent trajectory in a single denoising process. The continuous latent serves as the sole persistent generative state, from which tokens are read out at each step to scaffold subsequent latent updates. HC-DLM outperforms standard diffusion baselines on Sudoku, Countdown, and LM1B benchmarks.

  6. The Missing Primitive: Diagnosing and Repairing Structural Mathematical Reasoning in LLMs
    06arXivAI Research

    The Missing Primitive: Diagnosing and Repairing Structural Mathematical Reasoning in LLMs

    While LLMs perform well on math problems, their structural understanding remains questionable. This paper introduces 'Mathematical Primitives' and a new benchmark to evaluate reasoning across four dimensions: Discovery, Generation, Digestion, and Execution. Diagnostics show that raw accuracy masks specific capability gaps, and 'Discovery' is the primary bottleneck. To address this, the authors develop a primitive-privileged self-distillation framework that successfully transfers structured reasoning capabilities to student models, consistently boosting performance across benchmarks.

  7. Designing Unstructured Proteins: IDiom and RL-SAE Enable Interpretable Feature Control
    07arXivAI Research

    Designing Unstructured Proteins: IDiom and RL-SAE Enable Interpretable Feature Control

    Intrinsically disordered protein regions (IDRs) lack fixed 3D structures, rendering traditional structure-based design methods ineffective. To address this, researchers trained IDiom, an autoregressive language model, on a dataset of 54 million predicted IDR sequences (IDiom-DB). To control functions, they introduced RL-SAE (Reinforcement Learning with Sparse Autoencoder features), which rewards sequences activating specific target features. Across eight tasks, RL-SAE achieved a 90% target feature activation rate, outperforming activation steering at 24%.

  8. DMAD: Recasting Distribution Matching as Adversarial Distillation for Fast Visual Generation
    08arXivAI Research

    DMAD: Recasting Distribution Matching as Adversarial Distillation for Fast Visual Generation

    Traditional Distribution Matching Distillation (DMD) requires training an auxiliary diffusion model to track the student's shifting distribution, incurring massive computational costs. DMAD bypasses this by reformulating distribution matching as a classification task. Using two discriminator heads on a shared backbone to distinguish real and teacher samples from student samples, DMAD directly learns log-density ratios. At optimum, these linear losses mathematically recover the DMD gradient. DMAD achieves state-of-the-art few-step generation performance across image, video, and audio-video models, including SDXL and Wan2.1.

  9. Introducing Olmo-core 3: Open, Scalable Training Infrastructure for Trillion-Parameter MoEs
    09Hugging FaceLLM

    Introducing Olmo-core 3: Open, Scalable Training Infrastructure for Trillion-Parameter MoEs

    Olmo-core 3 redesigns the MoE training stack, transitioning from FSDP to a DDP-based system that keeps experts resident on GPUs. In benchmarks using NVIDIA B300 GPUs, a 47B MoE model achieved 52k tokens/sec per GPU, representing a 2.7x throughput increase over previous versions. Supporting MXFP8 precision and advanced parallelism, Olmo-core 3 scales up to trillion-parameter configurations, offering researchers an efficient, fully open-source infrastructure.

  10. AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures
    10Hugging FaceAI Agent

    AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures

    General LLMs often struggle with specific enterprise workflows and tools. ServiceNow's AutoSynthData addresses this by comparing a target model's failures with a stronger teacher's successes to produce "capability specification cards." It then generates brand-new, realistic training tasks. Each task undergoes rigorous positive and negative verification gates, with failed tasks repaired via a critic loop. Tested on EnterpriseOps Gym, this targeted synthetic data approach boosted the target model's Hybrid Pass@1 rate by 7.2 percentage points, representing a 35% relative improvement.

Past issues