Skip to content
The AI-First Web

Cheap AI checkers could make reviewing every agent action affordable

OpenAI's Decisions API and TypeSafe's Jev point to a cheaper way to watch AI agents. The evidence is still early.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
Cheap AI checkers could make reviewing every agent action affordable

AI-generated image for WebPulse. About our images

Key finding

TechCrunch figure, basis not stated: monitoring of that kind with Jev (Hugging Face scenario): $2.94 (Source: TechCrunch report on OpenAI's Decisions API and TypeSafe's Jev (September 30, 2026))

Oversight is only as good as what it costs to run. If watching an AI agent costs as much as the agent's work, most organisations will watch a sample, or nothing. A report from TechCrunch suggests that cost may be falling, though the evidence is early.

What was announced

At OpenAI's Dev Day on Tuesday, CEO Sam Altman mentioned a new "Decisions API" in an aside. TechCrunch reports that it seems to offer similar functionality to Jev, a model TypeSafe AI released earlier this month for software automation.

Altman's pitch was that developers can narrow OpenAI's Luna model to a single pick from a short list. The list might label a picture or select how an agent should behave. In his words: "By focusing the model on that choice, we can make it extremely fast."

The caveats matter. OpenAI released it as a limited preview. TechCrunch says it has not yet seen developers testing it, and it is not clear how closely it matches Jev. TypeSafe did not answer TechCrunch's questions.

How a decision model works

TechCrunch describes Jev as a "super-powered classifier" built on an LLM, a large language model. Developers give it a set of choices. It returns each choice as a probability, cheaply and at high speed. This describes Jev, and what Altman said about the Decisions API.

General background, not from the source: a chatbot builds free text one word at a time, and each word takes another pass through the model. A fixed list of options removes most of that work. The model only has to score a handful of choices. Less output generally means less time and less cost. The source does not explain how Jev or Luna reach their speeds.

TechCrunch's view is that LLMs as we know them are comparatively slow and expensive for a lot of software. It says developers using Jev alongside LLMs have found the result faster and cheaper. The source gives no measurements for that.

If TechCrunch's reporting holds, many software steps look like choices, where free text may be the wrong tool. That is an interpretation, not a tested finding.

The use that matters to leaders: watching agents

OpenAI has added security measures after a series of incidents in which its agents misbehaved on the open internet. One is a separate model that watches for bad actions, at what TechCrunch calls "significant compute cost."

Shapor Naghibzadeh, who has long worked in cybersecurity and runs the startup QueryStory, believes a model like Jev could bring that cost down sharply. He built a demo for last weekend's hackathon. It compares each agent action with the task the agent was given. Actions it is confident are bad get blocked. Uncertain ones are flagged for review. The rest go ahead.

$2.94
TechCrunch figure, basis not stated: monitoring of that kind with Jev (Hugging Face scenario)
Source: TechCrunch report on OpenAI's Decisions API and TypeSafe's Jev (September 30, 2026)
$372
TechCrunch figure, basis not stated: same monitoring with a frontier LLM (Hugging Face scenario)
Source: TechCrunch report on OpenAI's Decisions API and TypeSafe's Jev (September 30, 2026)

TechCrunch gives both figures in a sentence about stopping the Hugging Face incident, without a basis. It does not say how many actions were checked, what task was run, or which frontier model was used. Treat them as an unverified illustration of scale, not a price list.

TechCrunch says such monitoring could in theory have stopped that incident. The source gives no detail on what the incident was, so readers cannot judge the claim. It is a hypothetical, not a tested result.

Why the price changes the decision

Card networks do not review a sample of purchases by hand. They score every transaction automatically and send only the doubtful ones to people. The same pattern could apply to agents, but only if the scoring is cheap enough to run on every action.

TechCrunch calls Jev "arguably cheap enough to run on every agentic action." Review could become a default layer, if the cost and calibration claims hold.

What is still unproven

TechCrunch names the key question: how well calibrated each model's outputs are to real life. A checker that says 95% bad must be wrong about 5% of the time, not 30%. Poor calibration means too many wrongful blocks or too many missed attacks.

TypeSafe CEO Diogo Almeida says his company's advantage is the synthetic data it uses to train for statistically useful outputs. That is the company's claim. TechCrunch also notes that other startups offer similar models.

Questions to put to your team

First: which agent actions in our systems are checked today, and which are trusted by default? Second: what do we pay for that checking, and is it the reason coverage is partial?

Third: if a vendor offers a fast checker, how will we measure its calibration against our own tasks before we rely on it? Fourth: where does a flagged action go, and who is on call to decide?

Cheap checks do not make agents safe. They make it possible to ask, for every action, whether it matches the task. The remaining work is deciding who answers when the checker is unsure.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: TechCrunch.

Share this insight