Skip to content
The AI-First Web

An ITSM AI agent pilot centred on three narrow jobs, its builders report

Its builders say a four-person IT team's pilot argues for starting agents with one repeatable, easy-to-reverse task

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
An ITSM AI agent pilot centred on three narrow jobs, its builders report

Photo: Yan Krukau / Pexels

In brief
  • The authors of a VentureBeat guest post built an AI agent for ServiceNow and report that real pain points were narrower and more repetitive than their roadmap assumed.
  • Their pilot with a four-person team centred on catalog items, ticket trend analysis and license clean-up. Time and savings figures are the authors' own claims, not audited results.
  • Before funding an agent, pick one repeatable, measurable, reversible workflow and ask who approves each write.

It is tempting to plan an AI agent as an assistant for everything a platform does. A guest post on VentureBeat describes a team that did not set out to build a general agent, yet still found its roadmap too broad. The authors say the real problems were "narrower and more repetitive" than that roadmap assumed.

The idea worth keeping: the narrowest workflow is also the easiest to measure, approve and undo. Scope is not only a way to save money. It is the main control on an agent that can change your systems.

What the team built

The authors built a "Digital Worker" for ServiceNow, an IT service management platform. They describe their company's ITSM data as plentiful but hard to turn into action.

The post has two authors: Richard Mendis, a chief marketing and strategy officer, and Brian King, who works on AI products at Bytemethod.ai. The post describes the tool as the authors' own company's build. So the people assessing it are the people who made it.

They piloted the tool with their own four-person ITSM team on a production instance. Three jobs came out of it.

The first is building catalog items, the request forms and workflows employees use to ask IT for things. The second is analysing 60 to 90 days of incidents and requests for trends. The third is checking premium licenses to find ones that are paid for but unused.

About 20 seconds
Authors' reported time to generate a catalog item, before human review
Source: Mendis and King, VentureBeat guest post (October 4, 2026)
80%
Catalog development cost cut the authors expect
Source: Mendis and King, VentureBeat guest post (October 4, 2026)

The authors say the 20 seconds covers the workflow and associated scripts. A person then previews the result and finalizes it before it goes live.

How it works under the hood

The team first used "browser use" in early 2025. That means the AI clicks through an application's screens like a person. They later switched to Model Context Protocol (MCP), a standard that lets an AI call a system's APIs directly. The authors say this made tasks faster.

They also wrote their own coordinating software, which the post calls an agentic harness. It works with several language models. It pairs the model's probabilistic reasoning, which can vary from run to run, with deterministic code, which gives the same result every time, for validation and execution.

That pairing matters for a manager. Put simply, the model reasons, and fixed code validates and executes. That is our reading of the design. A system that only chats can be wrong without consequence. One that edits your platform cannot.

To teach the agent house rules, the team used reusable SKILL.md instruction files plus company knowledge and memory. It also reads their requirements documents and flags missing information before building anything.

Controls the authors say limit the damage

The authors describe four safeguards designed to keep people in charge. Before any write operation runs, a human must review and approve it. The agent connects through OAuth under the user's own ServiceNow identity, so role-based access limits apply. Update sets, ServiceNow's way of packaging configuration changes, give the team a familiar way to move those changes through deployment or roll them back. Finally, observability tracks the agent's actions, their success and their cost.

An agent with write access can multiply whatever it touches, including its mistakes. These controls are meant to keep that multiplier small. The agent is tied to the user's permissions, and update sets cover configuration changes. The post offers no evidence of how well the controls work in practice.

25%
Capacity the authors estimate they can reclaim
Source: Mendis and King, VentureBeat guest post (October 4, 2026)

Read the numbers with care

This is a first-person account from the builders. The 80% and 25% figures are what the authors "expect" if they apply the tool to all future catalog work. They are not measured results.

The pilot also involved one four-person team. The post gives no independent testing and no error rates. Treat it as one data point on how to scope an agent, not proof of savings.

The authors also say two parts were harder than expected. One was getting the agent to read the organization's requirements format reliably. The other was finding the real pain points.

What leaders should ask

Ask your team for one candidate task. It should come up constantly, follow set rules, produce results you can count, and be simple to undo. The authors advise starting there rather than with a broad agent. If no task fits, the roadmap is not ready.

Ask who approves every write, and whose permissions the agent uses. Ask whether each change can be rolled back, and whether you can see what each run did and cost.

Ask whether the savings are measured or projected. Then ask what happens to the saved hours. In this pilot, the stated goal was more time for higher-value platform work.

A broad agent is a bet on a roadmap. A narrow one is a test you can read. Start with the test.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: VentureBeat.

Share this insight