Skip to content
Innovation & Growth

AWS lets AI propose the fix for an outage; a human clicks to approve

AWS shows how to turn an incident diagnosis into a ready-made repair. The approval step now carries the risk.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
AWS lets AI propose the fix for an outage; a human clicks to approve
In brief
  • AWS published a design in which an AI agent prepares the fix for an outage and a human gives the final approval.
  • The approval click becomes the main safety control, and the person clicking now checks finished-looking work rather than investigating.
  • Ask who can add tools to the allowlist, what evidence sits beside the approve button, and whether approvers ever reject a proposal.

The on-call engineer used to do two jobs at 3 a.m.: work out what broke, then fix it. AWS has now published a design that hands most of the first job, and the preparation of the second, to software. The human keeps one task. They decide whether to press approve.

That is a real change in what an engineer is for. It is worth understanding before your team adopts something similar.

What AWS describes

AWS DevOps Agent is an AI agent that triages incidents using metrics, logs and application topology. It produces a root cause analysis and recommended actions. AWS notes that organisations usually keep such agents in an observe-and-report mode. The agent diagnoses but does not change production.

A new AWS blog post, published October 7, 2026, shows how to add a second stage that prepares the fix. The agent finishes an investigation and emits an event to Amazon EventBridge, AWS's event routing service. A rule triggers a small Lambda function, which fetches the investigation summary. That summary goes to a longer-running Lambda Durable Function.

The durable function runs a loop. It asks an Amazon Bedrock model what to do, runs a tool, feeds the result back, and repeats until the fix is complete.

How the safety controls work

Two controls limit what the model can do. First, the model may only pick from an allowlist of purpose-built tools. AWS gives examples such as reading a Lambda function's configuration or updating an IAM policy statement. IAM is the service that decides who can do what in an AWS account.

Second, tools are split into read-only and mutating. Read-only tools run on their own. A mutating tool makes the function stop and wait for a person.

The waiting is cheap. The function saves its progress, pauses without using compute, and resumes after the approval signal. AWS says it can pause for minutes, hours or days.

Up to 1 year
Longest run for a durable function
Source: AWS Machine Learning Blog (October 7, 2026)

The demo is deliberately small

AWS tests the design on a simple fault. A Lambda function times out because its limit is set to 3 seconds. The agent finds this cause. The model reads the configuration, confirms the 3 seconds, and proposes raising the limit to 30 seconds. The engineer enters an approval and the change is applied.

3 seconds
Original function timeout
Source: AWS Machine Learning Blog (October 7, 2026)
30 seconds
Timeout proposed by the model
Source: AWS Machine Learning Blog (October 7, 2026)

AWS says the post reduces mean time to resolution. The post reports no measured figures for that. It is a walkthrough of one simple case, not a result from production.

The approval click is now the security control

AWS is direct about this. Its post calls the human approval gate "the security control" and says the summary and the proposed fix are AI-generated. Whoever signs off must read the precise values the tool will apply, such as which function is touched and what the new timeout is, and judge for themselves that they are right.

Here is the lesson. When software prepares the fix, the human is no longer investigating. They are checking someone else's work, and that work arrives looking finished. A tidy proposal built on a wrong diagnosis is harder to doubt than a blank screen.

It is a familiar shift, as with autopilot in aviation. The skill that matters moves from doing the task to catching the machine's mistake. Catching mistakes at night, under pressure, is a different skill from fixing the fault yourself.

Two details in the design deserve attention. The tool examples include editing IAM policy, which is far more sensitive than a timeout. AWS also presents the design as easy to grow. Teaching the agent a new kind of repair means editing a registry of tools, while the orchestrating code stays as it is. That makes widening the machine's authority a small edit. The allowlist is therefore a governance decision, not a technical footnote.

Questions to put to your team

Before anyone builds this, ask these questions.

Which tools are on the allowlist, and who signs off when one is added? Treat each addition as a change in what the machine may do.

Does the approver see the exact values of the change, or only a summary? AWS says the full detail must be checked.

How does the approver judge whether the diagnosis behind the fix is right? What evidence sits next to the approve button?

Who is the approver at 3 a.m., and have they ever rejected a proposal? An approval gate that is never refused is not working as a control.

Faster fixes are worth having. But the engineer's judgment is the part of this system you cannot automate away, so build the workflow around protecting it.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: AWS.

Share this insight