A new research approach called "Image-to-Code" enables AI coding agents to generate executable Blender programs from single photographs, but the agents lack the ability to reliably assess the geometric accuracy of their own creations.

What Happened

The method, detailed in a study by Li et al., instructs a coding agent to write and run Blender code iteratively. The agent writes code, executes it, views the resulting 3D scene, and revises the code until the output matches the input image. Because the final product is code rather than a static model, users can explicitly inspect and modify objects, geometry, layout, and camera positions.

To evaluate performance, the researchers introduced LEGO-Bench, a dataset comprising 208 images from 104 indoor and outdoor scenes built with 443 registered assets. Unlike real photos, which lack precise 3D ground truth, or simple synthetic scenes, which appear unrealistic, LEGO-Bench uses professionally built simulator scenes rendered to look natural. This allows for automated scoring based on exact geometry, depth, and object assignments.

The benchmark assesses scenes on three axes: validity (whether a usable artifact was produced), reconstruction (geometric accuracy), and appearance (visual similarity via pixel-by-pixel comparison). Tests on six GPT configurations showed that while all models delivered working scenes almost every time, geometric accuracy varied significantly. GPT-6 Astra, the top performer, achieved 53.4 percent on indoor scenes and 39.6 percent on outdoor scenes, while weaker configurations scored around 15 percent.

Notably, increasing the reasoning budget significantly improved performance. On an office test subset, GPT-6 Astra’s score jumped from 32.3 percent to 61.8 percent when given more computational resources for reasoning, indicating that current limitations are partly tied to inference constraints rather than pure architectural failures.

Why It Matters

The study highlights a critical failure in current agentic workflows: unreliable self-assessment. Analysis of the agents' work steps revealed that models frequently made poor initial attempts, revised code in ways that undid previous progress, and failed to recognize when their scenes had degraded. In one instance, GPT-6 Astra’s accuracy dropped from 33.9 to 4.4 percent late in the refinement process because it failed to detect that its own edits had worsened the scene.

When asked to judge which of two versions better matched the original, the models’ geometric judgments landed near or below chance level. This indicates that current agents cannot trust their own visual evaluation of 3D geometry. In response, the authors developed LEGO-Plugin, an extension that replaces self-judgment with concrete measurements and shields correct progress from regressive edits. The plugin improved all six tested models, with weaker agents seeing gains up to 62.7 percent, while the top model saw a marginal improvement of about two percentage points.

The authors conclude that while current agents show promise, they are not yet accurate enough for high-fidelity reconstruction, noting that a big gap remains between a working code artifact and a faithful 3D reconstruction. This is evidenced by downstream vision tasks: when reconstructed scenes were used for object detection, segmentation, and depth estimation without extra training, performance remained unremarkable. Object detection reached roughly half the performance of the specialized DINO model, while segmentation and depth estimation lagged further behind specialized tools like SAM 3 and Depth Anything 3.

The industry landscape is already reacting to this capability gap. 3D software makers are gearing up for these kinds of agents, with Unity releasing official plugins for Claude Code and Codex. Meanwhile, alternative approaches are emerging that skip code generation entirely, such as the Atlas world model from World Labs, and Google DeepMind’s GenCeption, which uses a video model for depth estimation and segmentation that matches the performance of specialized models.

The Bottom Line

Current AI agents can generate executable 3D scenes from photos, and increasing reasoning budgets significantly boosts their performance. However, they cannot reliably verify their own geometric accuracy, necessitating measurement-based tools like LEGO-Plugin. While the field is advancing with industry support and alternative direct-reconstruction methods, a substantial gap remains between agent-generated scenes and specialized 3D models.