- An agent that writes and loads its own tools can fix itself, but it can also modify and delete, as AWS's Sandhya Subramani said.
- She argued that sandboxing the code-execution environment, constraining permissions and running evals are what make self-modifying agents safe to trust.
An agent that writes its own tools turns each unexpected request into new code that no one has reviewed. Sandhya Subramani, a senior developer advocate for GenAI at AWS, demoed this on the AI Engineer show. She called it powerful. She also argued that it needs sandboxing, evals (automated checks of agent behavior) and tight permissions before anyone trusts it in production.
What was said
Subramani showed an agent built with the open-source Strands Agents SDK. It had a system prompt and three tools: an editor, a shell and a tool loader. Its tools folder held no tool files.
She asked it to help with a complex math problem. It wrote a math calculator tool and used it straight away, with no restart. In a second demo, a travel planner built its own sub-agents.
She pitched this as self-healing. When something breaks in production, the agent could spot the bug and repair itself. She said her longer demos sometimes break when the audience pushes hard. In those cases the agent notices the error and tries to fix its own tool or agent.
Then she turned to the risk. "If this meta tooling or this meta agent can spin off tools and agents, it can also modify and delete," she said. She asked whether anyone wants it deleting data they worked hard to build.
Her answer was layers of checks. Evals should test whether the agent reached the goal, made up an answer, picked the right tool and used the right parameters. They should also check how sub-agents talk to each other. Guardrails start with sandboxing the environment where the generated code runs, not just the agent's container. They also include limits on tools and permissions, plus telemetry to watch it all.
Why it matters
Our reading: the usual review step disappears. In a normal system, a person reads new code before it ships. Here the agent writes it, loads it and runs it in the same moment.
A mistake in that loop can be a wrong answer, or it can be a deleted file. For teams that build or buy agent software, the practical questions are specific. Where does generated code run? What can it touch? What record shows what it wrote? Subramani's list works as a checklist to put to a vendor or an internal team.
The other side
Subramani works at AWS and was showing off an SDK her company builds. A live demo is not evidence of how these controls hold up in real incidents. She did not say how well the evals catch bad tool code, or whether anyone reviews what the agent writes.
She was open about the downside. She said her demos can break, and she voiced her own worry about self-evolving agents. Yet she called the technology incredible. She also noted that the Python version of Strands wrote the TypeScript version by itself. That shows capability, but the talk did not say what review that code received.
Written by the WebPulse Newsroom with AI assistance, and checked by our editorial review: every quotation was verified against the recording's transcript. How we use AI.
The conversation this talking point comes from
- AI Engineer: Agents That Write Their Own Tools at Runtime — Sandhya Subramani, AWS (2026-10-04)





