Aivora
arXivAI AgentIntermediate

VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness

VISTA:為多模態 AI 打造的「長視域」互動式視覺輔助框架

2 min read
VISTA: Empowering Multimodal Agents with a Long-Horizon Visual Harness
The 30-second version

Developed by a team including Kaiming He, VISTA is a general-purpose visual harness that grants multimodal models long-horizon vision. It records raw observations in a lossless visual memory, allowing agents to actively retrieve and reorganize visual inputs during active reasoning. Tested on ARC-AGI-3, VISTA boosted Claude Opus 5.0's score to a perfect 100.00, completing games using 57.4% fewer actions than novice humans, proving its high generalization potential across visual puzzles.

Key points

01

Long-Horizon Lossless Memory

VISTA preserves past visual observations in their raw, original form to prevent any information loss over long decision steps.

02

Active Retrieval & Reorganization

The agent can actively query its visual history and dynamically reorganize its inputs to aid reasoning.

03

Superior Action Efficiency

On ARC-AGI-3, the VISTA-equipped model completed tasks using 57.4% fewer actions than first-time human players.

04

Simple and Generalizable Design

VISTA's minimalistic design allows easy adaptation across multiple diverse visual interactive environments without heavy customization.

How it works

VISTA Interactive Visual Harness Workflow
Send raw framesQuery past statesInput reorganized viewsDecide next stepUpdate env stateInteractive EnvironmentLossless Visual MemoryActive Retrieval &ReorgMultimodal LLMReasoningExecute Action

Why it matters

While multimodal models struggle with memory in long-horizon tasks, VISTA shows that a simple, lightweight external harness can unleash their reasoning potential. Without retraining the foundation model, it achieves superhuman action efficiency on tough visual-reasoning tasks like ARC-AGI-3. This paves the way for highly adaptable multimodal agents in both virtual and real-world interactive domains.

Who it affects

  • AI Developer
  • AI Researcher
  • Product Manager

How to use it

  1. 1Interactive visual game and puzzle solving
  2. 2Autonomous multimodal agents requiring long-term visual state tracking
  3. 3Complex visual GUI or web environment execution and automated navigation

Limitations & caveats

  • Growing visual history may increase the load on the model's token context window over time.
  • Currently validated primarily on visual games and benchmarks; real-time, highly dynamic 3D physical environments may present unsolved challenges.

Related

AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures
Hugging FaceAI Agent

AutoSynthData: Generating Targeted Training Data from Enterprise Agent Failures

AutoSynthData:以企業 Agent 的失敗為師,自動生成高規格微調訓練資料

AutoSynthData is a framework by ServiceNow that analyzes an enterprise agent's failures against a stronger teacher to automatically generate and validate high-quality synthetic training data, bridging crucial performance gaps.

2 min read
KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux
arXivAI Agent

KaliBench: Evaluating and Boosting LLM Command Generation on Kali Linux

KaliBench:首個 Kali Linux 資安工具指令生成基準測試,助 8B 模型直逼 685B 巨獸

KaliBench is a fine-grained benchmark for evaluating LLM CLI command generation on Kali Linux. While open-weight models score below 42% accuracy, training an 8B model using KaliBench's runtime-free verifiable rewards allows it to rival a 685B MoE model.

2 min read