4DCodeBench: Benchmarking AI Agents on 4D Inverse Graphics and Dynamic Scene Code Generation
4DCodeBench:評估 AI Agent 動態場景 4D 反向圖形學與程式碼生成能力的全新基準
4DCodeBench introduces a novel benchmark for 4D inverse graphics via code generation. Agents observe videos of dynamic events (like fluids, deformation, and fracture) and reconstruct them into executable graphics programs. Benchmark results reveal that models excelling at static reconstruction still struggle with simulating complex physical dynamics.
Key points
4D Inverse Graphics
Evaluates agents' ability to translate dynamic video inputs into executable 3D graphics and physical simulation programs.
Diverse Physical Phenomena
Includes real-world and synthetic scenes spanning complex dynamics like deformation, fluid flow, and fracturing.
Static vs. Dynamic Gap
Experiments reveal that frontier models with strong static reconstruction capabilities still fail to reliably model complex dynamics.
Executable Abstractions
Requires agents to distill visual observations into high-level abstractions like program loops and physical simulation parameters.
How it works
Why it matters
Traditional inverse graphics focuses on static shapes, but 4DCodeBench expands the horizon to the spatiotemporal (4D) domain. This drives the development of agents with deeper physical world models, paving the way for automated 3D game development, virtual scene reconstruction, and physically grounded embodied AI systems.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
How to use it
- 1Automated Game & Simulation Asset Creation: Converting real-world physical videos directly into editable, interactive physics engine code.
- 2Embodied AI World Modeling: Allowing robots to predict and simulate complex physical environments via dynamic code generation.
Limitations & caveats
- Current frontier LLMs and agents exhibit low reliability and success rates when generating code for complex dynamic physics.
- The benchmark depends on executable simulation environments, which may not perfectly model or evaluate extremely complex, non-linear physical dynamics.
Related
Less Decoder is More Encoder: Extracting Robust 3D Geometric Representations via Novel View Synthesis
減少解碼器反而增強編碼器:從新視角合成中提煉強大三維幾何表徵
This paper reveals how expressive decoders dilute geometric learning in Novel View Synthesis, and proposes SNAP—a self-supervised framework that restricts decoders to force encoders to learn robust 3D representations.
What Should World Models Forget? Stratified Retention for Continual Adaptation
世界模型該遺忘什麼?以「分層保留」實現持續適應環境的能力
This paper argues that world models must not avoid all forgetting, proposing 'stratified retention' to distinguish permanent physical laws from dynamic, environment-specific facts that require timely revision.

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude
艾倫人工智慧研究所開源 AstaBrief:比 Claude 快 3.5 倍的 8B 科學報告生成模型
Allen Institute for AI (Ai2) has open-sourced AstaBrief 8B, a specialized model for scientific report generation that achieves a 3.5x speedup over proprietary pipelines while maintaining high citation accuracy.