Current leading AI agents can meet basic physical and semantic requirements for LEGO design but struggle to match human-level design quality, revealing gaps in spatial reasoning and constraint satisfaction under discrete choice problems.
BrickBench is a benchmark that tests AI agents' ability to design buildable LEGO sets from text descriptions. Agents must select parts from a library and satisfy physical constraints, semantic requirements, and design quality—combining reasoning about local details with global structure. The paper includes BrickAgent, an environment for agents to build and validate designs.