- In a University of Maryland and AWS test of six GPT configurations, agents judged pairs of 3D scenes near or below chance on geometry.
- Delivery and accuracy are different measures. The top model scored 53.4 percent indoors and 39.6 percent outdoors on the LEGO-Bench accuracy score.
- Ask what independently measures your agents' output. In this study, a plugin combining measurements with other fixes lifted weaker models by up to 62.7 percent.
An agent that reviews its own work is only as reliable as its judgment of that work. A new test of GPT-based agents on 3D scene building shows that judgment can be weak. The lesson is that an AI agent should be checked against measurements, not against its own opinion.
What the researchers built
A team from the University of Maryland and AWS created LEGO-Anything, a method they call Image-to-Code. As The Decoder reports, a coding agent gets a single photo and writes code for Blender, a widely used 3D program. It runs the code, inspects the outcome and keeps revising until the scene resembles the photo.
The output is a program, not a picture. It spells out the objects, geometry, layout and camera position. A person can read, edit and test it like any other code.
To score the agents, the team created LEGO-Bench. Its 208 images come from 104 indoor and outdoor scenes, built on 443 registered assets. The images are rendered from professionally built simulator scenes. They look natural, but the exact geometry stays hidden and acts as an answer key.
Working output is not correct output
All six tested GPT configurations delivered a usable scene almost every time. Accuracy was another matter. The strongest model tested, GPT-6 Astra, scored 53.4 percent indoors and 39.6 percent outdoors on the benchmark. Weaker configurations scored around 15 percent.
These are LEGO-Bench accuracy scores, not simple success rates. The source does not define how the aggregate score is calculated.
This gap matters for anyone who reads a status report. "Task complete" tells you a file exists. It does not tell you the file is right. In this test, delivery was nearly universal and fidelity was not.
The agents could not judge their own geometry
The researchers went through the agents' step-by-step work to see why accuracy suffered. Three issues came up most often: weak opening drafts, edits that wiped out earlier gains, and self-assessment that could not be trusted.
The last one stands out. Shown two versions of a scene, a model had to choose which matched the original better. Its geometric judgments landed near or below chance. The team's takeaway: an outside measurement, not the agent's own opinion, should decide whether a change helped.
Think of a writer proofreading their own draft. The same blind spot that produced the error hides it on the second read. A loop of "try, look, revise" only helps if the "look" step is reliable. Here it was not.
A plugin with outside measurements lifted scores
The team's answer was LEGO-Plugin, an add-on that needs no extra training. It does three things. It anchors the first scene in the reference image. It replaces the agent's self-judgment with concrete measurements. It also shields correct progress from later edits that make things worse.
The plugin improved all six models. Weaker agents gained most, with boosts of up to 62.7 percent. The top model, already far ahead, moved up by roughly two percentage points.
The source reports one combined result. It does not say how much of the gain came from the measurements, how much from the image anchoring and how much from the edit protection. It also does not say whether 62.7 percent is relative or in points.
Spending more on thinking also helped. When the reasoning budget rose, Astra's score on an office test subset climbed from 32.3 to 61.8 percent.
Limits of the finding
This is one benchmark on one task, built from simulator scenes. It tested only six GPT configurations, so it says nothing about agents from other vendors. It does not show how agents judge their own work in finance, support or software. It does show that fluent iteration can coexist with poor self-assessment.
A large gap also remains. The authors say a big distance separates a working result from a faithful reconstruction. Used as input for other vision tasks, the scenes reached roughly half the performance of the specialized DINO model on object detection. The gap was larger for segmentation and depth estimation against SAM 3 and Depth Anything 3. The authors say the scenes show promise but are not yet accurate enough.
Questions for your team
If your organisation uses agents that revise their own output, ask four things.
First, what does the agent check its work against? A second opinion from the same model is not a measurement.
Second, who owns that check? If no person or tool outside the agent scores the result, the agent is grading itself.
Third, is the reported success rate about delivery or about accuracy? In this study the two numbers were far apart.
Fourth, would a weaker model with strong external checks be enough? In this benchmark, the weaker models gained the most from the plugin. The study does not say which part drove that gain, and it did not compare cost.
An agent that cannot tell whether its last edit helped is not a reviewer. In this test, the ruler had to sit outside the agent.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: The Decoder.





