- Stack Overflow's engineering blog argues that an LLM system fails by acting wrongly at scale, not by crashing, so it needs controls a normal service does not.
- Its proposal is one gateway for all model traffic and one decision ID. Cost limits, failover and the kill switch all hang off that gateway.
- Ask your team whether they can stop one customer's or one feature's AI in seconds, and whether they have tested it.
When a normal web service breaks, it says so. It returns an error, a dashboard turns red and someone is paged. An AI system that breaks often keeps answering. It just answers wrongly, and it can take wrong actions at scale, and fast, before anyone looks.
Stack Overflow's engineering blog makes this point in Part 5 of a series on running large language models (LLMs) in production. The post says a traditional service that fails returns a 500 error. An LLM system that fails "can take wrong actions at scale, fast." The post treats that gap as the reason operability is not optional.
Control is a place, not a policy
The post's central idea is small. Send all model traffic through one chokepoint. Give every decision one ID. Keep the business logic unaware of which vendor answered.
The lesson here is that control over AI is a single place in your architecture, not a written policy. Think of a building's fuse box. You can cap, measure and cut power at one panel instead of walking every room. Companies that let each team call a model directly have no such panel.
What the single gateway does
The post describes one gateway that every model call passes through. It is the one place where several jobs get done.
First, it measures spend. Cost becomes visible per customer, per feature and per model, so the team can see which feature or oversized prompt is driving the bill. Second, it enforces limits. The post says spend and rate caps must use shared state across servers, because each server checking its own memory would allow its own full share. It also calls for a platform-wide ceiling.
Third, it routes and fails over. The post advises checking a circuit breaker before calling a provider. A circuit breaker stops calls to a service that is down, so failures take milliseconds instead of 60-second timeouts. It also advises one overall deadline instead of a fixed timeout per model. Without that, total wait time grows with every fallback.
Fourth, it holds the kill switch. The post wants a runtime switch that works per customer and per capability, and that lands across all servers in seconds. It says a 30-second local cache is not "seconds." It also describes a middle setting, HUMAN_ONLY, where the system keeps proposing but stops acting on its own. If a server cannot read the switch, it should assume the cautious mode.
The post also sets example targets for a system that makes decisions. One is that at least 99.9% of decisions return a proposal. The post's reasoning is that failing to decide is the real outage, not an error code. This is one example target from the post, not an industry standard, and it is not the same as ordinary uptime.
The bill is unmonitored, not high
The post says LLM cost stays hidden until the invoice arrives. It grows with factors teams rarely watch, such as the size of each call, the number of calls behind one request, retries, and a wordy prompt that someone added. Two multipliers get special mention. Retry storms happen when a flaky check keeps re-calling the model. Unbounded tool loops happen when an agent keeps going.
In the post's words, "LLM cost isn't fundamentally high; it's fundamentally unmonitored." The fix it offers is structural, not a cheaper model. Measure at the gateway and cap what can run away.
The post also says that once a cheaper model passes your tests for a task, routing that task to it makes calls 5–20× cheaper. That is the author's description of what is common, not a measured result. Treat it as a reason to test, not a forecast.
Slow is rarely the model's fault
Leaders often assume a slow AI pipeline needs a faster model. The post offers an example that points elsewhere. A pipeline that matches 10,000 items against another 10,000 seems tied to the model. Profiling, which measures where time actually goes, tells a different story. Most of the time, 90%, goes to one nested loop that makes 100,000,000 comparisons.
The post pairs this with a rule: lock a quality baseline before any speed work, and revert any change that makes answers worse. Fast and subtly wrong is a bad trade.
Keep the vendor out of the business logic
The post argues that providers change every few months, so vendor code should not leak into core logic. It recommends ports and adapters. The business logic talks to an interface your team defines, and one adapter per provider talks to the vendor. A build rule fails if domain code imports a vendor library. In the post's telling, "can we switch models?" then becomes an afternoon's work, not a project.
The same discipline covers fallbacks. A fallback model is a different model, so it needs the same tests. The post warns that if a backup model quietly lowers answer quality, the result is a worse outage than the one the backup was meant to cover.
Questions to put to your team
This is one publisher's engineering guidance, not a study. It offers no incident data. Still, its checklist translates into plain questions.
Can we stop AI for one customer or one feature in seconds, without a deploy? Have we tested that switch under real failure? Can finance see AI spend by customer and by feature before the invoice arrives? Is there a platform-wide spending ceiling? Do we record which model made each decision? Has our fallback model passed the same tests as the primary? When a provider goes down, does work route to a person?
An AI system you cannot see, cap or stop is not automated. It is unattended.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Stack Overflow.





