Skip to content
The AI-First Web Talking point

Because models stay probabilistic, agents need durability built into the harness

Temporal's Melanie Warrick argued that probabilistic models need structure that keeps state and pauses for humans.

W
WebPulse Newsroom
AI-assisted · 2 min read
Share on X LinkedIn
Because models stay probabilistic, agents need durability built into the harness
In brief
  • Melanie Warrick of Temporal said AI models are probabilistic and hard to make reliable, so teams add structure around them, called a harness.
  • She argued that durability should be part of that harness, so an agent that waits for a human can survive a crash without losing its place.

AI models give uncertain answers, and agents run them in loops with little supervision. Melanie Warrick, who works in developer relations at Temporal, told the AI Engineer show that these models are probabilistic and hard to make reliable. Teams therefore add structure around them, which the field calls the harness. She argued that part of that structure should be durability: the system survives a crash and resumes where it stopped. She added that the field is also trying to improve the models themselves.

What was said

Warrick described an agent as a loop: it thinks, acts and perceives, then repeats. Her demo ran three such loops for an ice cream delivery service, covering fleet, customer and dispatch. A human could step in to approve a decision.

That human step is where things break, she said. Code it as a plain function call and the agent can stall while it waits. The request can also vanish if the service goes down. People are slow compared with software. As she put it, "humans are not going to respond in 200 milliseconds."

Her answer was Temporal's tooling. A workflow records each step the agent takes. Model calls and tool calls run as separate activities. Two primitives handle the human: a wait condition pauses the work, and a signal from the person resumes it. The show notes say she killed the worker mid-approval and restarted it with no lost work. The mechanism, in her words: "your event log is getting replayed it's not redone."

Why it matters

Our reading: the risk in agents is less the single wrong answer than the loop around it. A probabilistic step inside an unattended loop runs again and again. If the system also loses its place after a crash, a small flaw can grow. On this view, reliability is being engineered around the model rather than inside it. That is WebPulse's argument, not Warrick's.

For anyone building or buying agent software, her talk suggests concrete questions. Where is the agent's state stored? What happens when a server dies mid-task? How long can it wait for a person? She also gave a test for involving a human: "the cost of being wrong is high."

The other side

Warrick has a stake here. Temporal is free and open source, and the company charges for managing state on customers' behalf. She also said she is not an expert on the harness. She noted that people debate whether the model, tools, memory and guardrails belong in it.

Durability keeps work from being lost. She did not claim it makes the model's choices correct. The human check has its own weakness, which she named: "alert fatigue is real," and people end up approving everything. She called the question of when to involve a person hard and case by case, with no formula.

Written by the WebPulse Newsroom with AI assistance, and checked by our editorial review: every quotation was verified against the recording's transcript. How we use AI.

The conversation this talking point comes from

Share this insight