Aivora
arXivAI ResearchIntermediate

WorldAuditBench: Evaluating Multimodal Agent Capabilities in Interactive 3D World Auditing

WorldAuditBench:首個評估多模態 Agent 3D 虛擬環境審計能力的全新基準測試

2 min read
WorldAuditBench: Evaluating Multimodal Agent Capabilities in Interactive 3D World Auditing
The 30-second version

As 3D simulations grow vital for testing intelligent behavior, identifying environmental anomalies (e.g., floating assets or traversable walls) is key. The authors introduce WorldAuditBench, featuring 213 auditing tasks across 13 Unreal Engine 5 and Three.js environments. Evaluating five frontier models under two paradigms (VLA-led exploration and end-to-end VLM reasoning), the study shows AI models achieve success rates of only 6.6% to 42.3%, lagging significantly behind the human benchmark of 83.4%.

Key points

01

Coupling Action and Reasoning

3D world auditing demands that agents coordinate navigation actions to search environments systematically while using visual reasoning to identify anomalies.

02

Rich Benchmark Design

Comprises 213 diverse anomaly tasks across 13 environments built on Unreal Engine 5 and Three.js, spanning 5 distinct anomaly families.

03

Two Paradigms Evaluated

Tests two workflows: VLA-based exploration followed by VLM anomaly detection, and an end-to-end VLM agent where reasoning directly steers actions.

04

Significant Human-AI Gap

Frontier models struggle with success rates of just 6.6% to 42.3%, severely underperforming compared to the human baseline of 83.4%.

How it works

Comparison of Auditing Approaches in WorldAuditBench
兩階段架構 (VLA + VLM)端到端 VLM Agent人類測試者 (Human)
NavigationVLA 動作模型探索場景VLM 視覺推理直接引導動作直覺式主動探勘
Anomaly Detection探勘完後由 VLM 統一分析探勘過程中即時整合分析即時感官與邏輯判斷
Best Success Rate6.6% - 42.3% (併計)6.6% - 42.3% (併計)83.4%

Why it matters

This research addresses a core bottleneck for multimodal agents: gathering and interpreting physical evidence in 3D spaces. It has immediate practical implications for automating QA in game development, metaverse building, and simulation prep for Embodied AI, reducing manual debugging effort.

Who it affects

  • AI Developer
  • AI Researcher
  • Content Creator

How to use it

  1. 1Automated QA for game engines and virtual assets to catch physics clipping and rendering bugs.
  2. 2Auditing simulator fidelity and environmental consistency before deploying reinforcement learning models.

Limitations & caveats

  • Evaluations are restricted by a fixed exploration budget, which may bound performance in highly complex 3D environments.
  • The benchmark utilizes 13 environments built with UE5 and Three.js, which might not represent all commercial rendering pipeline anomalies.

Related

Fixing the "Timing Shortcut": A Breakthrough in Non-Invasive Brain-to-Text Decoding
arXivAI Research

Fixing the "Timing Shortcut": A Breakthrough in Non-Invasive Brain-to-Text Decoding

排除「時間捷徑」漏洞:非侵入式腦機介面解碼技術的新突破

Researchers revealed that recent breakthroughs in non-invasive brain-to-text decoding relied on a "timing shortcut" of word durations rather than actual brain signals. Their SimpleB2T method eliminates this shortcut, slashing the word error rate to 36.6%.

2 min read