There's a statistic making the rounds that ought to embarrass our whole industry: even among AI proofs-of-concept that <em>succeed technically</em>, only about one in twenty produces any positive outcome in the real world. The models work. The demos impress. And then... nothing. The pilot ends, the tool sits there, and eighteen months later nobody remembers why it existed.
In healthcare, where the promised payoffs — lower costs, less clinician burnout, better care — are desperately needed, this pattern isn't just wasteful. It's a clue. And I've become convinced the diagnosis isn't bad models. It's a missing <em>system</em>.
First, untangle the question
When people talk about "AI in healthcare," they're usually blending a half-dozen different problems into one buzzword: make clinicians' lives easier; get clinicians to trust algorithmic output; cut administrative waste; deliver the best care cost-efficiently, because no one has infinite dollars; and beneath all of it, take better care of actual patients — which, never forget, is the goal every clinician will drive themselves into the ground pursuing.
The discipline I keep coming back to is embarrassingly simple: <strong>what question are we asking, and can our data answer it?</strong> Most dead pilots died at conception, because they were an answer wandering around looking for a question. A model that "predicts X" is not a question. <em>Who acts on that prediction, what do they do differently, and how would we know it helped?</em> — that's a question. The 5% that survive tend to be the ones that could answer it on day one.
The loop, not the tool
But even a well-posed pilot dies if it's built as a <em>tool</em> rather than as part of a <em>loop</em>. Here's the vision that I think deserves to be the organizing idea for this whole era — it's been developing in health policy circles for a quarter century, and the technology has finally caught up to it: the <strong>learning health system</strong>. My working definition: an integrated health system that harnesses data and analytics to <em>learn from every patient and every clinician's practice</em>, then feeds what it learns back to clinicians, patients, and everyone else in cycles of continuous improvement.
Read that twice, because every word is load-bearing. <em>Every patient</em> — not the narrow slice enrolled in formal trials. <em>Feeds back</em> — the knowledge returns to the people delivering and receiving care, rather than accumulating in a dashboard nobody opens. <em>Cycles</em> — it never finishes.
Let me make it concrete with the best real-world example I know of, currently underway in a health system serving vast rural communities. The problem: cancer patients often don't receive guideline-recommended care — some because of access or insurance, some because the guidelines genuinely don't fit them ("the guideline says this, but it won't work for you because of these other conditions"). And the guidelines themselves are built on evidence everyone acknowledges is thinner than we'd like. Two intertwined failures: patients missing the standard, and a standard that fails patients.
The team's design attacks both at once. An AI tool gathers everything relevant about a patient and lays it beside what the guidelines recommend — synthesis work no busy clinic has time to do by hand. That synthesis goes to the clinician <em>and patient together</em>, who discuss it and decide — sometimes per the guideline, sometimes deviating for good reason. Then the crucial parts: the decision <em>and its rationale</em> get captured; the team measures whether the process actually helped both parties feel better informed; and then — the step most pilots never reach — they follow what happened to those patients. Survival, quality of life, everything. And over a horizon of years, those outcomes flow back to the people who write the guidelines, so the standard itself gets better and more patient-specific.
Notice what the AI is in this design: one component in a loop that includes clinicians, patients, outcome tracking, and guideline revision. Ask the question, gather the data, decide with humans, capture outcomes, improve the standard, repeat. Nobody expects one team in one health system to settle it alone — the point of pilot-then-scale is that the <em>model of working</em> propagates.
Pilot, test, scale — and never expect perfect
That phrase — pilot and scale — is engineering thinking, and healthcare needs much more of it. The engineering method in one breath: you can't just sit and think — build something, test it small, expand what works, and accept that <em>nothing is ever perfect</em>: you improve continuously until you hit an inflection point that changes everything, and then you start over from fundamentals. AI in medicine is exactly such an inflection point — unsettling and exciting in equal measure — which makes the discipline more necessary, not less.
It also exposes why the tool-not-loop failure mode is so lethal now. A CT scanner changes maybe once a decade; you learn it, trust it, use it. These new methods evolve <em>constantly</em> — which means a health system without a working feedback loop isn't just missing improvement; it can't even tell when its deployed tools have drifted out of date. The learning cycle isn't a nice-to-have on top of the AI. It's the only safe container for AI this fast-moving. (And the volume problem cuts the other way too: no clinician alive can read every relevant journal article anymore and be sure they're delivering the best care. Keeping up now <em>requires</em> system-level learning.)
The saddest sentence in innovation
"It was a wonderful tool, and it just sat there." I've come to think there's a single question that predicts, at design time, whether that sentence will someday be written about your project: <strong>when this tool is wrong, or the world changes, what feeds back, to whom, and what do they change?</strong> If the answer is silence, the tool is already dead; the pilot just hasn't ended yet. Wonderful ideas don't die because they were bad. They die because there was no feedback loop to improve, integrate, and carry them forward — no loop that would have let them live.
Build the loop first. The tool is the easy part.





