Skip to content
The AI-First Web Talking point

When agents run the query, the human skill left is judging the answer

A Composio engineer used Datadog daily for six months without ever opening its dashboard, and froze when she had to.

W
WebPulse Newsroom
AI-assisted · 2 min read
Share on X LinkedIn
When agents run the query, the human skill left is judging the answer
In brief
  • Sarah Simionescu used Datadog daily for six months without opening its dashboard, then froze when a coworker asked her to check an alert.
  • When an agent plans, queries and summarises, the person's job shifts to judging the answer, which raises questions about training.

Sarah Simionescu, a member of technical staff at Composio, argued on the AI Engineer show that the dashboard is dead. Agents now run the queries, so people no longer need to learn each tool. Her own story shows the catch: she used Datadog daily for six months and never opened its dashboard.

What was said

A coworker once asked her to look at a Datadog alert. Datadog is a software-monitoring tool. She had to log in, then froze and could not find her way around. She said this was no knock on Datadog: "I had genuinely never opened the dashboard."

She traced the problem to 2022, when one bug report meant five tools and five query languages. Next came AI, where, she said, an LLM would write "a sometimes correct query to get you the answer you need."

Her demo showed how an agent builds an answer. Claude first asks a Composio search step for the tools it needs. It named three tasks: fetch Slack messages, search Sentry issues, search Datadog logs. Composio returned the tools plus a plan, such as finding the Slack channel ID before querying messages. The agent then pulled Datadog and Sentry data in parallel and scanned the code. A pull request with a fix was up in under five minutes.

She said MCP, a standard for connecting agents to apps, is not enough alone. Agents get no map across apps, and too many tools overwhelm the model.

Why it matters

Simionescu spoke about interfaces. Our reading is that her story also shows a change in the work. When an agent plans, queries and summarises, the person's job is to judge the result.

Separately, she said the dashboards she once opened daily are becoming more unfamiliar to her. In her demo, the answer rested on a plan, parallel pulls and saved results the person did not write. Judging that takes some knowledge of the systems behind it.

Managers can ask who on the team could check an agent's work against the raw tools. Training is the second question. If agents do the daily tool use, newcomers may never gain that familiarity.

The other side

Simionescu did not present this as a problem. She called the dashboard's death great news for most of the room, and said her freeze felt like a preview of what is coming.

In the demos described, runs succeeded and no step for a person to check them was mentioned. Her results were early and unreleased, and Composio builds the interface she promoted. The excerpts leave open how people keep judging systems they rarely touch.

Written by the WebPulse Newsroom with AI assistance, and checked by our editorial review: every quotation was verified against the recording's transcript. How we use AI.

The conversation this talking point comes from

Share this insight