Models that fully completed the course: 1 of 4 (Source: DrivingBench report (September 2026))
The brake was a person and a few lines of code
In a parking lot, at speeds capped between 1 and 8 mph, a general-purpose chatbot has now steered a real car through a cone course. Only one of four models finished. The more useful question for executives is what kept the test safe.
The report's safeguards were all external: speed limits written into code and a human ready to brake. The model refusals it describes were an obstacle the team worked around, not a safety finding. This shows a pattern worth noting. The protection the team relied on was the one they built around the model.
This matters beyond cars. Companies are giving AI agents tools that act on real systems. The limits you can audit are the ones enforced outside the model.
What DrivingBench did
DrivingBench is three Bay Area tech workers, according to 404 Media. Four models took the wheel, in effect: GPT-6 Astra, Claude Fable 5.1, Grok and GPT-5.6 Sol. The car was a 2022 Toyota Corolla.
The team wanted to know whether general-purpose models can drive in the real world. They laid out a cone course in a large parking lot. It had a left turn, a right turn, a straight, bends and a parking zone as the finish.
Each model got up to three attempts in the same chat. After a failed attempt, the team sent a generic prompt asking the model to reflect on its mistakes. Then the model tried again.
GPT-6 Astra finished on its second attempt, in about five minutes. Claude Fable 5.1's third attempt and Astra's first attempt each got about halfway. All other attempts stopped before the first corner.
The team says most failures were errors of sight. At the first diagonal row of cones, the models misread which side the lane was on.
How a chatbot steers a car
The setup used comma four, an add-on device that connects to the car's internal network. It sends camera frames and telemetry to a laptop. The laptop runs the chatbot through an MCP server, a standard way to give AI tools it can call.
The model had three tools. One returns camera frames plus speed and steering. One sets a motion command. One brakes immediately.
A motion command takes a speed, a duration and a steering percent. The team mapped 100% to 180 degrees of steering-wheel angle. Their example: 60% becomes 108 degrees, which openpilot, comma's driving software, turns into a tire angle of about 7.8 degrees.
The car keeps moving while the model thinks. A new command replaces the old one. When a command expires, the car brakes.
That design turns slow thinking into a physical risk. In one Fable attempt, the car drove for only 31 seconds of 190. Most of the time went to reasoning while it sat braked.
Refusals were an obstacle, not a control
The team reports that some models, especially GPT-6 Astra, sometimes refused to drive. They cited safety, even in an empty lot with speed caps and a human ready to brake.
Calling the task a simulation did not hold. In some trials the model saw real images and realized it was real.
The team says what worked best, for some reason, was renaming the MCP tool "DrivingBench Sandbox." That change came alongside a revised prompt. Models drove consistently only after both. The team does not say which change mattered, or why.
So the cause is unknown. This is one experiment with one setup. It does not show that a label changes a model's judgment. It is a caution: when a small change in wording coincides with a model agreeing to act, that behavior is hard to audit.
Limits sat outside the model
The team's real safeguards were external. The tool and controller capped speed at 0.5 to 3.5 meters per second, about 1 to 8 mph. A separate emergency stop cancelled motion above 6 meters per second, about 13 mph.
They left openpilot's driver monitoring on. An operator sat in the car with a foot above the brake throughout. The steering wheel also cannot turn faster than 100 degrees per second in this setup.
Those limits held regardless of what any model decided. That is the design lesson.
Learning in the chat, with caveats
Astra and Fable showed signs of learning from their own reflections. After its first try, Astra's notes promised shorter moves at low speed near bends. On its second try its speed stayed at or under 0.8 m/s, and it completed the course.
The caveats are the team's own. Each model was tested once. Attempts share one chat, so they are not independent. The setup was one Corolla and one comma device. Some prompt details were off: the steering example overstated how long a 90-degree turn would take.
Questions to ask before an agent touches a real system
If your teams are wiring AI agents to physical or financial systems, ask these questions.
First, which limits are enforced in code or hardware, and which depend on the model behaving? Second, can a renamed tool or a reworded prompt change what the agent will do? Test it, and change one thing at a time.
Third, who is the human with a foot on the brake, and can they intervene fast enough? In this test, a command kept running while the model thought. Fourth, what does one successful demo not tell you? Here it was one model, one car and one lot.
The lesson is that capable agents make external limits more important. The model can help with the driving. It should not be the only thing deciding when to stop.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: DrivingBench.





