Skip to content
The AI-First Web

Meta test: a separate controller helped an AI agent use a larger budget

On ProgramBench with GPT-5.5, the controller setup rose from 64.1% to 71.5%. The baseline used about 18% of its calls at the top setting

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
Meta test: a separate controller helped an AI agent use a larger budget
In brief
  • Meta Superintelligence Labs researchers tested a separate controller for AI agents. On ProgramBench with GPT-5.5, it improved from 64.1% to 71.5% as the call budget tripled.
  • Their baseline, with no separate controller, stayed near 64% and used about 18% of its available calls at the highest setting. This is one comparison.
  • Teams should measure how agents spend compute, not only how much they are given. All results come from the researchers' own tests.

A bigger budget is not a better plan

Hand a team more money and you expect better work. Anyone who has managed a project knows it often fails to happen. The money needs someone deciding what to fund, what to drop and when to stop.

New research from Meta Superintelligence Labs applies that idea to AI agents. VentureBeat reported the work on October 9. In one test, a larger allowance of compute did not help the baseline agent. The agent with a separate controller turned it into better scores. For anyone paying for agents, the question is how an agent spends its budget, not only how big the budget is.

What the researchers built

The researchers call the problem metacognitive control. They define it as "assessing one's own progress and using that assessment to decide what to do next."

Many agents make this call in the same step as the work itself. They pick their next action from a growing pile of past attempts, tool outputs and errors. Useful findings can get buried in that pile.

Meta's answer is a harness called the Meta-Reasoning Agent. It splits the job in two. Workers do the tasks. A controller decides what happens next. Workers can be single model calls or coding agents that inspect files, change code and run tests.

The controller runs a four-step loop. It assesses new results. It proposes possible next actions. It evaluates those actions against the remaining budget. Then it dispatches workers, or stops and submits the result it has.

Think of a site manager who walks the floor, reads the reports and assigns the next crew. The crews do not decide among themselves.

Why the extra budget mattered

The key detail is memory. Every worker output and controller note is stored as an artifact with its own identifier. The controller keeps a short summary of the task and pulls up detail only when needed.

The system also tracks which past results feed into which new tasks. A test report might prompt a repair, which another worker then checks. Those traces add up to a map of where the agent put its effort.

71.5%
ProgramBench hidden-test pass rate, Meta-Reasoning Agent with GPT-5.5
Source: Meta Superintelligence Labs researchers, as reported by VentureBeat (October 9, 2026)

Codex scored 58.0% on the same benchmark in the researchers' evaluation. ProgramBench asks an agent to rebuild a program from its documentation and executable references.

The more telling result is what happened as the budget grew. The researchers raised the ProgramBench allowance from 400 to 1,200 model calls. This budget comparison was run with GPT-5.5 on that one benchmark.

64.1% to 71.5%
Meta-reasoning score on ProgramBench, 400 to 1,200 calls (GPT-5.5)
Source: Meta Superintelligence Labs researchers, as reported by VentureBeat (October 9, 2026)

Direct control stayed near 64% over the same range. Direct control is the researchers' main baseline. It is a variant of the same harness in which one agent makes control decisions over its own history.

About 18%
Share of available calls used by direct control at the highest setting
Source: Meta Superintelligence Labs researchers, as reported by VentureBeat (October 9, 2026)

Put plainly, the baseline used about 18% of its calls at the highest budget. The controller version's score went from 64.1% to 71.5% over the same increase. The source gives those two points, not a full curve.

What this means for people who pay for agents

The lesson here is that agent spending deserves the same oversight as any other spending. The size of the allowance tells you little. What counts is whether the agent turns it into progress.

Two other details matter for governance. In the ProgramBench setup, the controller can check worker claims by inspecting the Git repository through a read-only interface. And the artifact graph gives a way to ask whether an agent used its compute productively or produced disconnected attempts.

A record that shows how work built on earlier work is also something a reviewer can read.

Limits to keep in view

These are the researchers' own results on four benchmarks, with three frontier models. At the largest budgets, meta-reasoning scored higher in all 12 matched comparisons. The budget-scaling result above comes from a single comparison on ProgramBench with GPT-5.5.

The approach is not free. The control cycle adds overhead and extra model calls. At modest budgets, the baseline sometimes came out ahead. The framework took the lead only once more compute was available.

The researchers did not ship runnable code with the study. The paper describes the controller and worker prompts, the memory design and the tools. Developers can use that description to try their own version.

Questions to put to your team

Ask whether your agent evaluations measure how compute is spent, or only the final score. Ask what share of an agent's allowed calls it actually uses, and why it stops.

Ask whether intermediate work is recorded in a form someone can audit. Ask how the system behaves at small budgets, where added control may cost more than it returns.

More compute is a purchase. Knowing where to spend it is the part you still have to build.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: VentureBeat.

Share this insight