Skip to content
Innovation & Growth Talking point

A small finance model trained for under $500 in compute beat its 235B sibling

Snorkel AI's Charles Dickens described a small model trained on expert-shaped, human-reviewed data.

W
WebPulse Newsroom
AI-assisted · 2 min read
Share on X LinkedIn
A small finance model trained for under $500 in compute beat its 235B sibling
In brief
  • A 4 billion parameter finance model, trained for under $500 in compute, beat its 235 billion parameter sibling on 290 held-out questions.
  • In our reading, a pass/fail training signal is only as good as its reference answers, so checking data may be the scarce skill.

A small AI model trained for under $500 beat its much larger sibling on realistic finance questions. Charles Dickens, a research scientist at Snorkel AI, explained how on the AI Engineer show. His case: for business work, reliability and specialization can outweigh raw size. The small model had 4 billion parameters (adjustable settings). The big one had 235 billion.

What was said

Dickens said the team took annual reports (10-K filings) from the SEC's EDGAR system. A 30 billion parameter model turned them into about 6,900 database tables and wrote one question and answer per table. Financial experts helped develop the question types.

Verification had three layers: programmatic checks that table and column names were real, reviews by independent AI agents, and manual human review. The team held out 290 questions as the test.

The agent answered by calling tools that query those tables. Dickens named three failure modes, even in frontier models. They assumed tables or columns existed. They flooded their own context with careless queries such as select star. They recovered badly from errors. The 235 billion model would "repeat the same failed strategy instead of considering the error messages and adapting," he said.

Training used a pass/fail reward from a judge model checking reference answers. Compute cost about $420 and the judge about $40. Simple setups won: single-table data gave the most lift, and one binary reward beat an expert-built rubric. In his words: "The bottleneck was never reasoning depth. It was just tool use".

Why it matters

Our reading: the pass/fail reward trained tool discipline, such as checking schemas and recovering from errors. That reward only teaches good habits if the reference answers are right. Dickens did not test this link, so it is our inference. He did say small models can be trained with reinforcement learning (trial-and-error training) when quality data represents the target tasks.

For buyers, a sharper vendor question follows: who verified the data, and how? Dickens also stressed that in financial and legal work you must be able to audit the agent.

The other side

Dickens called this a first pass. The test was 290 questions, each with one checkable answer. The win was over the same model family's larger version, not every large model.

A model generated the question and answer pairs. The excerpts do not say how many humans reviewed. No experiment isolated the effect of expert verification. The $500 covers compute and judge fees, not data building or expert time, so the true cost for a team is unknown.

Written by the WebPulse Newsroom with AI assistance, and checked by our editorial review: every quotation was verified against the recording's transcript. How we use AI.

The conversation this talking point comes from

Share this insight