Skip to content
The AI-First Web

Postman and AWS say missing context, more than missing tools, broke its AI agent

An 11-year-old product was built for human eyes. Making it readable to an agent exposed hidden assumptions.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
Postman and AWS say missing context, more than missing tools, broke its AI agent

Photo: ThisIsEngineering / Pexels

In brief
  • Postman and AWS report that missing context caused more agent failures than missing capabilities, after Postman adapted an 11-year-old, interface-driven product for an AI agent.
  • The lesson is that a product's screen carries knowledge nobody wrote down. An agent cannot read it, so someone has to make it explicit.
  • Leaders should ask which tools and what context their agents see, and which actions need human approval.

The screen was doing a job nobody wrote down

A new employee learns an office by walking around. They see which door leads where and which folder holds what. Long-time staff forget they ever learned this, because the building itself carries the knowledge.

Software works the same way. Postman, a popular tool for building and testing APIs, found this out while building Agent Mode, its AI assistant. The company and AWS describe the work in a post on the AWS Machine Learning Blog, published October 9, 2026.

The team expected model quality and prompt wording to be the hard parts. They were not. The harder problem was that the product had grown over 11 years around people clicking through screens. An agent does not click through screens. It reasons over data.

This shows a shift that executives should plan for. Making a mature product usable by AI is mostly not a model problem. It is the work of writing down what the interface used to imply.

Why more tools made the agent worse

An agent acts through tools, which are small functions it can call. Postman first built many tiny ones, such as opening a request or updating one field. Each step had to go back to the model before the next could start. Users watched the agent plod through what they saw as one action.

A bigger problem followed. In Postman's testing, tool-selection errors rose once the agent could see more than about 40 tools. It called tools that did not exist, passed wrong arguments, or picked a plausible but wrong tool. Larger or newer models reduced this, the report says, but did not remove it.

~40 tools
Tool count where selection errors rose
Source: Postman and AWS, AWS Machine Learning Blog (October 9, 2026)

The current design searches a database of tool descriptions. It narrows more than 170 tools to about 15 for the request. It then hands those to a separate sub-agent, so the model sees only what the task needs.

170+ to ~15
Tools narrowed per request
Source: Postman and AWS, AWS Machine Learning Blog (October 9, 2026)

Tools tied to what was open on screen

Many of Postman's internal functions depended on interface state. Some needed a certain tab to be open. Others opened new tabs as a side effect. To read a request, the agent first had to bring up its tab. That mirrored a person's clicks rather than using the underlying data.

Postman is now separating tools from tabs. Agent Mode can send a request in the background without an open tab, though the user must still approve it.

For its API Catalog, the team replaced several narrow read tools with one query tool. The agent sees the layout of the underlying ClickHouse database tables and writes its own queries. The engineering job changes from building a tool per question to modelling the data well once.

Context beat capability

Postman first assumed missing tools would be the main blocker. The report says missing or incomplete context caused more failures than missing capabilities. Tools still mattered, as the tool-count problem above shows. Context here means where the user is, which items are active, and what has already happened.

The obvious fix did not work. Feeding the model the product's existing data objects produced little use, because those objects were shaped for display and data transfer, not reasoning. Postman built a dedicated handler for each type of item, and each one distills what the agent needs.

User-written content, such as descriptions and API specifications, can also fill the model's working space. Clutter hides the useful facts well before the model runs out of room, the report warns. Postman treats that space as a budget to be managed.

The controls around the model

Agent Mode asks for user approval before actions that change application state. Enterprise admins can switch on Amazon Bedrock Guardrails, which strips personal data before it reaches the model. The report notes that production testing and monitoring are still needed.

On the model side, Postman runs Anthropic Claude models through Amazon Bedrock. It can route different jobs to different models. It can choose cross-Region profiles that either maximise capacity or keep processing within a defined geography. It caches stable parts of the prompt for one hour and changing parts for five minutes.

Postman has set zero data retention for supported Agent Mode models. Availability varies by model. Teams should confirm the setting in AWS's current data-protection documentation for each model they run in production.

What to take from one vendor's account

This is one company's experience, written by the company and its cloud provider. The post gives no figures on cost savings, speed or error rates. The 40-tool threshold comes from Postman's own testing and may differ elsewhere.

The pattern is still useful to test in your own products. If an agent will work inside a mature application, the interface assumptions are part of the project.

Questions for your team

Ask how many tools each agent can see at once, and who decides. Ask which of your product's functions only work because a screen is open. Ask what context the agent receives, and whether it is shaped for reasoning or copied from what the screen shows.

Ask which actions need human approval, and where personal data is removed. Ask whether the data-retention setting was checked for the specific model in use.

The screen was never only a display. It held the product's memory. Before an agent can use your product, someone has to say that memory out loud.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: AWS.

Share this insight