Skip to content
The AI-First Web

Keep the AI to one step and make the rest of the system ordinary code

A Stack Overflow blog post argues safe agents only propose. A separate, plain program acts after approval.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
Keep the AI to one step and make the rest of the system ordinary code
In brief
  • A Stack Overflow blog post argues that AI agents in high-stakes systems should only propose decisions, while separate, tested code applies them after approval.
  • The model is confined to one step, its output is checked against a fixed structure, and every decision is written to an append-only record.
  • Leaders can ask their teams whether any agent writes to business records directly, and what happens when a model's output fails validation.

Much of the demo-driven talk about agents centres on capability: loops, tool calls, models that "figure it out." A Stack Overflow blog post argues for a different first step: make the agent boring.

In a post published October 7, 2026, Stack Overflow's blog describes Level 1 of a six-level model for running language-model systems in production. It calls this level the determinism layer. The goal is to confine the model to a single step, so everything around it is ordinary code that can be tested.

The idea: an agent should advise, not act

Think of the old bank rule where one person drafts and another approves. The post builds a similar separation into software. In its design, an agent is a function that takes context and returns a proposed decision. It has no authority to change anything.

A separate, simple component, which the post calls the substrate, applies proposals. It does so only after the required approval. The post says this turns "an AI did something we can't explain" into "an AI suggested something, and here's exactly what approved it."

The practical gain is containment. If an agent is tricked or simply faulty, the worst it can hand over is a poor suggestion. The damage ends at the next checkpoint, whether that is an automated guardrail or a person who declines.

6
Levels in Stack Overflow's maturity model for LLM systems in production
Source: Stack Overflow blog (October 7, 2026)

How the design works

The post rejects free-form loops where a model decides its own next move. Each capability instead runs as a fixed sequence of steps, called nodes. The language model gets only the steps that call for judgment on messy input. Plain code handles the rest.

Every node has the same shape: it takes a typed state object and returns an updated one. That uniformity lets a team test each node alone. Only the model node is nondeterministic, meaning its output can vary between runs.

Eight nodes are shared across every capability. A new capability needs about four of its own. One of those is a verification node: deterministic, capability-specific code that checks what the model returned.

8
Nodes shared across every capability in the post's graph
Source: Stack Overflow blog (October 7, 2026)

Why free text is the weak point

The post says much of the fragility in LLM systems comes from letting the model answer in prose and then parsing it. Phrasing drifts between calls. Downstream code breaks on an unexpected sentence.

The fix is to force the model to return a structure that can be validated. Anything else counts as a failed call and is retried, with the validation error fed back. Retries are capped, and the system fails closed when they run out.

The post adds a caveat. Structure constrains form, not correctness. A well-formed answer with 0.99 confidence can still be wrong. Structured output removes parsing failures, not judgment failures. That is why the post treats it as the floor, with evals, verification and confidence checks above it.

Sending doubt to people

The post warns against taking confidence straight from the model, calling it miscalibrated. It composes confidence from three signals: the model's own signal, the verification result and a sampled judge.

A router then applies a single threshold and picks one of four paths. The post advises starting conservative, with everything going to a human. The threshold is lowered for a given slice of work only as data shows it is safe.

Every decision is then written as one immutable row. The post says the record should hold hashes of inputs rather than the inputs themselves. It should be append-only, with corrections superseding old rows. A ledger that can be edited, the post argues, cannot count as an audit trail.

Where loops are allowed

The post does not ban open-ended agent loops. It treats them as a rare, rail-guarded exception. A loop fits only when the number and order of steps depend on what is found along the way.

4
Rails the post requires on any agent loop: iteration cap, tool allow-list, full trace, same output checks
Source: Stack Overflow blog (October 7, 2026)

The post is direct about the temptation. A loop added "for flexibility" is usually a fixed graph in disguise. Flexibility a team does not need only makes the system harder to debug.

What the post does not show

This is design guidance from one blog post. The text gives no measured results from real deployments. It covers Level 1 of six, so evals and fuller confidence calibration are left to later parts. The routing described here is a first pass.

The post also names a cost: you give up some cleverness. In return, it says, you can reason about what the system will do. It states that in high-consequence systems such as money movement, healthcare and infrastructure, "it usually works" is a non-starter.

How that trade weighs for your own systems is a call for each team. The question is how costly a wrong action would be.

Questions to put to your team

Can any agent write to business records directly, or can it only propose? Which single step uses the model, and what checks its output? What happens when that output fails validation: a capped retry, then a stop?

Who sets the threshold for human review, and did it start with everything going to a person? Can we rebuild any past decision from a record nobody can edit? Where an agent loops, what are its step cap and tool list?

The lesson here is that trust in an AI agent comes less from the model than from the job description around it. An agent you can predict is one you can sign for.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Stack Overflow.

Share this insight