BrickBench: Evaluating Agentic Text-Conditioned LEGO Design
BrickBench:評測 AI Agent 進行文字導向積木設計的全新基準
This paper introduces BrickBench, a benchmark paired with the BrickAgent coding environment, designed to evaluate AI agents on text-conditioned LEGO set synthesis. Agents must select parts from a discrete library while jointly reasoning over local and global constraints. Evaluated across three settings of varying scales and part availability, results show that leading agents can successfully satisfy verifiable physical and semantic requirements, yet still lag behind human design quality.
Key points
Task & Benchmark Objective
Evaluates agents on generating assemblies from text prompts using discrete parts while satisfying semantic and physical constraints.
BrickAgent Interactive Environment
Provides an environment for coding agents to incrementally construct, inspect, and validate their designs.
Three-Tier Evaluation Settings
Evaluates agents across three settings varying in scale and part availability across validity, alignment, and design.
Gap with Human Designers
Leading agents meet formal physical and semantic requirements, but fall short of human craftsmanship and aesthetics.
How it works
Why it matters
Expanding generative AI into physical 3D structural reasoning is a crucial step for spatial and embodied AI. By framing discrete combinatorial assembly with physical constraints into a benchmark, BrickBench drives progress in LLM agent reasoning for computer-aided design, robotic assembly, and structural engineering.
Who it affects
- AI Researcher
- AI Developer
- Designer
- Student & Learner
How to use it
- 1AI-assisted CAD and 3D structural design
- 2Robotic autonomous assembly and physical construction planning
- 3Interactive educational and creative block model generation
Limitations & caveats
- Current leading agents still lag behind human professional designers in terms of overall design quality and aesthetic.
- The discrete part library and verification models may not fully capture all complex real-world physical dynamics.
Related

Building Reliable Data Analytics Agents: Lessons from NVIDIA's KDD Cup Harness Design
NVIDIA 於 KDD Cup 的實務經驗:打造可靠資料分析 Agent 的 Harness 設計原則
NVIDIA's KGMON team secured second place in KDD Cup 2026 by showing that with a small fixed LLM, optimizing the agent harness through unified data interfaces, constrained tools, and trajectory inspection dramatically improves analytics reliability.
RECAST: Active Evidence Construction via Adaptive Routing and Computation
RECAST:透過自適應證據路由,為 LLM 計算出正確的上下文
RECAST is an agentic framework that treats evidence construction as a sequential decision process, utilizing a Router-Compiler loop to actively compute and derive the optimal context for RAG rather than relying on passive retrieval.

Microsoft Releases Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Real-Harness Training
微軟開源 Agent Lightning v1.0:僅 3,500 行程式碼,直接用生產環境 Harness 訓練 Agent 的強化學習框架
Microsoft Research Asia has introduced Agent Lightning v1.0, a lightweight open-source framework that trains AI agents using their actual deployment harnesses via a transparent LLM proxy.