- Google's managed agents run both the agent loop and the sandbox on Google's servers, so developers skip building infrastructure.
- How much a team can see of a bad run may depend on what the platform exposes; the excerpts we reviewed did not cover replay, audit logs or liability.
Philipp Schmid of Google DeepMind, speaking on the AI Engineer show, presented a way to run AI agents where the vendor hosts everything. The decision loop and the sandbox, an isolated cloud computer where the agent runs code, both sit on Google's servers. That makes agents easier to ship. It also means a bad run happens on the vendor's machines. The excerpts we saw did not address who pays when an agent errs.
What was said
Schmid described a managed agent in the Gemini API. When a developer sends one request, Google's backend starts a cloud sandbox. There the agent can run code, write files, install dependencies, install the developer's own tools and call the developer's own APIs. Google manages the setup, so the developer runs no infrastructure.
The environments persist. A follow-up request can reuse the same one, and several agents can share it through its file system.
He contrasted this with the usual approach. Normally the developer's own code must catch each function call, run it somewhere and send the result back. Here, a speaker on the show said, "it's really just a single API call, and everything runs inside that remote sandbox." Teams can share finished agents with customers, the speaker added, "without writing any Terraform, without thinking about Kubernetes, microservices, Firecracker, or anything else."
The demo did show progress. The environment starts, the model thinks, then it installs Python dependencies and reads its skills. Files the agent creates can be downloaded or fetched through an API.
Why it matters
Our reading: when the loop runs on Google's servers, a team's view of a bad run depends on what the platform exposes. The excerpts showed step-by-step progress and file download. They did not address replay or audit logs.
For anyone building or buying agents, that raises practical questions. Can you replay a bad run? What records can you see? What does the contract say about harm done by an agent acting for you? The excerpts we have do not answer these.
The other side
Schmid was pitching convenience, not liability, and our excerpts are partial. A stretch of the talk is missing from them. He may have covered logging, limits or responsibility there or elsewhere. We cannot say he left them out, only that the excerpts we reviewed show no discussion of them. Contracts and terms of service, which we did not examine, may also settle who pays.
The design has a safety argument too. The show notes say an agent should not run on your laptop, and a separate sandbox keeps its actions away from your own machine. Teams that build their own agent infrastructure can get it wrong as well.
The developer still chooses which tools, sources and APIs the agent can reach. That choice shapes how much harm a mistake can do. Who pays when it does remains an open question.
Written by the WebPulse Newsroom with AI assistance, and checked by our editorial review: every quotation was verified against the recording's transcript. How we use AI.
The conversation this talking point comes from
- AI Engineer: Why AI Agents Should Have Their Own Sandbox — Philipp Schmid, Google DeepMind (2026-10-07)





