- BootLoops 1.0 is an open-source toolkit that lets AI agents run exact scientific calculations, paired with written protocols that define when a result counts as checked.
- WebPulse argues the lasting idea is that AI output needs a test it can fail. Schwartz warns that Claude declares victory early.
- Before trusting any AI-produced number, ask which independent route reproduces it and whether the check was shown able to fail.
Every organisation that lets AI produce numbers faces one question. How do you know the number is right? A new open-source project, BootLoops, offers one answer. It pairs its scientific tools with written protocols that define when a result counts as checked.
What BootLoops is
BootLoops 1.0 is an open-source toolkit that lets AI models run exact, high-precision scientific calculations. The repository describes a large software package alongside working protocols meant to make the results checkable. Most of the repository is the tools themselves, with install notes, engine lists and credits.
The choice of model does not matter, the repository says. You download the code, start whichever AI agent you use inside it, and the agent can then call the tools.
Claude wrote the code under the supervision of Matthew D. Schwartz, who maintains it. It uses the MIT licence. Maintenance falls to Schwartz, not Anthropic, which offers no support or fixes for it.
How the checking works
Every tool comes with a guide written for an AI reader. It covers the tool's purpose, the situations it suits, how to read its results, and the acceptance test its output has to clear.
The protocols live in a separate repository as plain-markdown files that most agent frameworks can read. They define what "done" means. One asks that a result be confirmed by a separate method, at test points held out of any fitting. The check itself must be shown capable of failing. Another asks the agent to recover a known answer before it touches real data.
Two more rules stand out. No source that fed a fit may later certify the result. And when a search for a mathematical relationship finds nothing, the agent must refuse rather than invent one.
The toolkit also handles error with care. Where the problem allows, it carries a proven error range through each step instead of a statistical one. Where it does not, it labels the error bars as statistical.
What the work produced
The Decoder relays Schwartz's account from an Anthropic guest post. In three months, he says, the group's output reached 36 manuscripts. They covered 18 fields and involved 19 co-authors. These are his figures, and we have not checked them.
In particle physics, Schwartz says, Claude computed 30 integrals within weeks, according to The Decoder's report. He says fifteen reproduced known results and fifteen were computed for the first time.
Reproducing known results resembles the planted-truth control the protocols describe, where an agent recovers a known answer before touching real data. It is not the same check. The account does not say how the split was designed.
Schwartz also says the results often became valuable only once domain experts set the direction. In ecology, for example, James O'Dwyer helped turn a finding into a better predictive model.
Why the tests matter
Schwartz is candid about the model's weaknesses. He says Claude declares victory too early, and that "done, with one asterisk" often means "not done at all." Automated checks are not reliable. Conclusions can be wrong even when the calculations are correct.
This is our interpretation: the protocols read like a response to that list. The argument here is that the scarce asset is no longer the answer. It is the test the answer must pass. The source does not demonstrate this. It is our reading of how the project is built.
Finance learned a version of this long ago. Auditors exist because the person who keeps the ledger should not be the only one to vouch for it. The protocols apply a similar separation to a model. The source that produced a number should not be the one to certify it.
The limits the authors state
The repository is open about its edges. Its self-test covers all 49 packages and takes a few minutes on a laptop. Most come back green. Three halt with a named error, because the data they need is not included. macOS is not part of release testing.
The authors tell users to validate outputs themselves. They say the tools are not intended for decisions in clinical care, actuarial work, payments, regulation or public safety.
The security note matters most to executives. Many tools evaluate the contents of their input files, so a file from someone else can run arbitrary commands. The authors say the integrity checks guard against accidents and are not a security boundary. Any agent that opens files from outside your team belongs in a sandbox.
Questions to put to your team
First, for any AI-produced number you rely on, which independent route reproduces it? Second, has the check ever been shown to fail on a case where it should? Third, did any data that tuned the answer also certify it? Fourth, will the system say "no relation found" or will it produce a plausible one?
Fifth, where do agents run, and what files can they open? Ask for the answer in writing.
BootLoops is one project and one author's account, so it proves no trend. It does show what a serious answer can look like. An AI that can compute is useful. An AI whose results arrive with a test they could have failed is something you can begin to evaluate. The authors still tell users to validate the results themselves.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: BootLoops-ai.





