- On a finance test, Charles Dickens of Snorkel AI said even frontier models showed three failures: invented schemas, self-flooded context and repeated failed strategies.
- A speaker on the show said reliability and specialization can outweigh raw model size for enterprise AI work.
On a finance test, even frontier models showed three recurring failure modes, said Charles Dickens, a senior applied research scientist at Snorkel AI. He spoke on the AI Engineer show. A speaker on the show put the cause as tool use, not reasoning depth.
What was said
Dickens's team built a question set called FinQA from SEC 10-K filings, the annual reports public companies file. It draws on roughly 7,000 tables. Each question has one verifiable answer. On this data, Dickens said, even frontier models showed three repeating gaps.
The first is schema hallucination. A schema is the layout of a database. The model assumes tables and column names exist, which Dickens said may come from what it saw in training.
The second is context flooding. The model fills its own working memory with badly planned requests, such as a query that pulls back every column. If it succeeds, the result can overwhelm the model's limits.
The third is poor recovery from errors. The 235-billion-parameter Qwen3 model would, a speaker on the show said, "repeat the same failed strategy instead of considering the error messages and adapting."
Dickens said the same failures appeared in insurance underwriting work. His team then trained a 4-billion-parameter model for under $500, using a simple pass/fail reward. A speaker on the show summed up the lesson: "The bottleneck was never reasoning depth. It was just tool use".
Why it matters
Our reading: these three gaps are old weaknesses that agents make faster. A person who guesses a column name hits an error and stops to think. An agent can guess, fail and retry at machine speed.
For buyers, that shifts the question. Ask how an agent handles a failed query, not only how large its model is. A speaker on the show said that for enterprise work, reliability and specialization can outweigh raw size.
A speaker on the show also asked whether an agent can be reliably audited, which matters in financial and legal settings.
The other side
The evidence is narrow. It comes from one finance dataset, plus a mention of insurance work. Dickens called the dataset a first pass.
The excerpts give no rates for how often each failure occurs. They do not show how any model behaves in a live enterprise system, or without human oversight.
The training reward was also decided by another model, GPT-5-nano, acting as judge. That is a practical choice, but it means the scoring rests on a second model's verdict.
Finally, the excerpts do not say whether the small model's training removed invented schemas. They say only that it improved results on this task.
Written by the WebPulse Newsroom with AI assistance, and checked by our editorial review: every quotation was verified against the recording's transcript. How we use AI.
The conversation this talking point comes from
- AI Engineer: Small Models, Big Results: Training a Finance Agent for Under $500 — Charles Dickens, Snorkel AI (2026-10-10)





