- For AI retrieval work, databases now return ranked, nearest-match results beside exact ones, so a poor answer can look like a normal result instead of an error.
- Code, access controls and audits written for exact answers may treat ranked results as exact, because the interface looks the same and nothing marks a poor match.
- Leaders should map where ranked answers enter, move permissions into the database, measure quality with evals and show users when nothing fits.
Quiet bugs in exact software
On September 29, Rails shipped version 8.1.4. A number of its repairs addressed defects that returned incorrect output while raising no alarm. Some queries could skip the condition that restricts results to one kind of record. A signed-record lookup came back empty for tables keyed on more than one column. Time-span calculations lost fractions of a second. Nothing crashed. Each defect handed back an answer that looked fine.
These are ordinary bugs in ordinary, exact software. They have nothing to do with AI. They matter here because they show the failure engineers find hardest to catch. An outage pages someone. A wrong answer flows quietly into an invoice, a report or a decision. In exact systems, such bugs are defects to be found and fixed.
AI brings a different kind of wrongness, and it is not a defect. For AI-style retrieval, a second contract now sits beside the exact one: ranked, nearest-match results. Many who sign off on AI systems may be accepting a probably-right answer without having chosen it. The larger risk is not the model. It is code and controls written for the first contract, which expect a query to return the truth or fail.
Two contracts, one interface
Sailesh Krishnamurthy, a VP of Engineering at Google, told Software Engineering Daily that he jokes a database has a single job. It must keep what it stores and return precisely what was requested. He dated that promise to the original SQL paper, from 1974. He was clear that this core purpose has not changed.
What changed is the workload. AI applications want customer records combined with documents, emails and notes. Krishnamurthy said that blend makes databases look more like search, where relevance and ranking decide what comes back. In his words, "now you have a mindset shift producing exact results to starting to produce inexact results." He added that databases are no longer one size fits all.
Here is how the second contract works, in plain terms. The system turns each piece of text into a list of numbers, called an embedding, that places similar meanings close together. A question becomes numbers too. The database then returns the handful of items nearest to it. Fast vector indexes, such as the HNSW method Krishnamurthy mentioned, search the likeliest neighbourhood rather than every item. They trade a little accuracy for speed.
The crucial point follows. There is always a nearest item. Unless someone builds one, there is no signal that says nothing fits. An exact query fails loudly, with an error, an empty page or a timeout. A ranked query fails quietly, because a poor match looks just like a good one. The interface stays the same: a query in, results out. So code written for exact answers will often treat a ranked result as exact, with nothing to tell it otherwise.
Permissions written for human queries
Access control is the next weak seam. In many applications, who sees what lives in filters developers added to each query. That holds while humans write the queries. It breaks when an agent writes its own SQL through an account that can read everything. Krishnamurthy cautioned against letting such an all-access agent answer on behalf of any user.
His fix moves the rule into the data layer. The agent may write any query, but the database shows only what the user may see: "So the query can have whatever it wants, but the view is on the fly, only restricted to the user's credential."
WebPulse reporting shows the alternative. At BlueHat Asia 2026, a researcher showed how Copilot in SQL Server Management Studio let a database owner gain sysadmin rights. Its read-only safeguard was a text-pattern check, and it had no separate restricted connection. Simple query tricks got write statements past it. In the final demonstration, the researcher's own login was granted the sysadmin role. Microsoft rates the flaw, CVE-2026-65669, critical. The guardrail lived in the tool, not in the database's own permissions.
On The Cognitive Revolution, Pete Johnson, who holds the Field CTO of AI role at MongoDB, broadened the argument. Poor data quality and weak security, he said, "don't get solved by AI, they get amplified by AI."
Look-alikes at machine speed
Models add their own blind spot. Hammad Bashir, CTO of Chroma, described his company's Context Rot research on Software Engineering Daily. It tested how models cope as more material fills their working memory, the context window. Models struggled to separate facts about similar but different things, such as the country of Georgia and the US state. More distractors made it worse.
Bashir stressed that the tests were contrived and mostly synthetic. Some models cope far better than others. Still, he said every model suffers from it to some degree. He sketched real versions: two colleagues with the same name, or two authentication systems in one codebase.
Then add speed. Agents, Bashir said, "don't query at a human rate. They query at an agentic rate." By his account, a person might search twice a minute at the very most. An agent may search every second or two, splitting each request into five or six parallel lookups. That is a description, not a benchmark. But the arithmetic is clear. At that pace, a repeated confusion can spread widely before any person looks.
Checking is possible, and the checks must be measured too. One team, reported by WebPulse, built a tool for AI answers that credit the wrong source. In one test it caught 138 of 139 such errors. It caught all 50 deliberately swapped sources. Its F1 score, a combined measure of hits and false alarms, was 0.846.
On a harder task, naming the exact correct source, it succeeded 50.3% of the time. The two figures measure different jobs. Spotting that an answer is wrong is easier than pinpointing what is right.
The accountants got here first
This shift has a precedent. Bookkeeping once rested on tie-outs, entries that balanced to the cent. As financial reporting leaned on estimates, auditors learned a new craft. They judged whether a range was reasonable, tested the assumptions behind it and demanded disclosure of uncertainty. Nobody treated an estimate as a ledger balance.
Software needs the same move, and the interface gives no prompt to make it. Code written for exact answers will often handle a ranked result as if it were a ledger balance. Exact systems already show what quiet wrongness costs, with no AI involved. In September, Next.js fixed seven flaws, and four involved caches serving the wrong content. In one, unpublished preview content could reach ordinary visitors who had not signed in. These are exact-contract bugs. A cache exists to return what was stored, and when it does not, nothing signals the mismatch. Ranked retrieval makes that silence part of the design.
The strongest objection comes from Krishnamurthy himself: "I think we have to embrace the chaos." The world is messy, and relevance unlocks value exact queries never could. His own example, flagging and blocking a fraudulent transaction before money moves, shows the upside. Demanding exactness everywhere would forfeit it.
But the objection does not defeat the argument. Krishnamurthy pairs the chaos with "being able to have evals to understand the quality and the business outcomes." That is the auditor's move: accept the range, then measure it. The problem is not inexactness. It is inexactness nobody chose and nobody measures, feeding systems built to assume the opposite.
Choose inexactness on purpose
First, map where ranked answers enter. Every point where a retrieval result feeds code that expects one correct record is a seam to inspect.
Second, move permissions into the database. Tie views and row-level rules to the user's credential. Do not rely on filters in application code or instructions in a prompt.
Third, measure before launch. Johnson argues that a team without existing metrics for a problem cannot tell whether AI improved it. Use the numbers the business already trusts.
Fourth, make uncertainty visible. Treat sources, confidence and the absence of a good match as real outcomes. A customer deserves a signal, not a confident wrong answer with no error message.
The executives and engineers who approve these systems will carry the blame when an answer was only probably right. They should at least have chosen it. The error message was the database's most honest feature. Do not retire it by accident.
Written by the WebPulse Newsroom with AI assistance, and checked by our editorial review: every quotation was verified against the recording's transcript. How we use AI.
Conversations this essay draws on
- Software Engineering Daily: Inside Google’s Database Infrastructure for the AI Era (2026-09-15)
- Software Engineering Daily: Chroma and Agentic Retrieval (2026-09-24)
- The Cognitive Revolution: Write, Change, Recall, Forget: MongoDB's Pete Johnson on How Retrieval Drives Agent Performance (2026-09-01)





