FrontierFinance Puts Every Major AI Model to the Test
Samaya AI just dropped FrontierFinance — and it changes how we measure AI on Wall Street. With 220 expert-curated queries and over 11,500 rubrics spanning the full investor workflow, this open benchmark is the largest and hardest finance AI test to date. And the results? Even the best models barely scrape past 50%.
Here is how the top contenders stack up — and why the verdict might surprise you.
Benchmark Results at a Glance
FrontierFinance covers six critical finance use cases: screening and discovery, company research, sector and macro analysis, financial data extraction, coverage and catalyst monitoring, and earnings and events. Every rubric is crafted by finance experts and verified through a multi-stage audit process. Systems are scored by how many rubrics they satisfy, with three independent judge models casting majority votes.
The overall rubric qualification rates tell a clear story:
- Samaya System (in-house): 50.8% — state-of-the-art, at roughly $0.80 per query
- Claude Fable 5 (Finance Agent v2): 49.2% — best frontier model, but at ~$3.20 per query
- Claude Opus 4.8 (Finance Agent v2): 45.0% — solid contender, half a step behind
- GPT-5.5 (Finance Agent v2): 43.5% — competitive but not dominant
- DeepSeek V4 Pro (Finance Agent v2): 40.5% — open-weight standout at 26% of the cost
- GLM 5.2, Kimi K2.6, Gemini 3.1 Pro: trailing the pack
What Makes FrontierFinance Different?
Older benchmarks like FinanceBench (150 queries) and BigFinanceBench (50 queries) mostly test financial data extraction — pulling numbers from SEC filings, calculating ratios, answering "what is the revenue?" type questions. FrontierFinance goes much deeper. Its queries demand synthesis across qualitative and quantitative data, causal analysis, forward-looking judgment, and even formatting requirements with specific presentation rules.
By difficulty scoring, FrontierFinance is significantly harder than any public alternative. Its mean Bradley-Terry hardness score sits at 2.3, well above FinanceBench's median. The hardest queries — screening and discovery, sector and macro analysis — leave every system scoring below 35%.
The Quality-Cost Tradeoff Nobody Is Talking About
Here is where the comparative review gets interesting. FrontierFinance plots every system on a cost-versus-quality scatter, and the picture is stark:
- Claude Fable 5 delivers top-tier quality on the Finance Agent v2 harness, but at over $3 per query, scaling it across thousands of daily investment tasks gets expensive fast.
- DeepSeek V4 Pro delivers 40.5% quality at just $0.83 per query — roughly the same absolute cost as Samaya's in-house system, but 10 percentage points lower on quality.
- Gemini 3.1 Pro on the web search harness costs under $0.50 but only hits around 25%, making it cheap but unreliable for serious finance work.
Verdict: If raw quality is your only metric, Claude Fable 5 wins among frontier models — but it comes at a premium. If value per dollar matters, DeepSeek V4 Pro is the budget king. And if you can use a specialized system, Samaya's in-house agent offers the best overall tradeoff: highest absolute score at a fraction of the inference cost.
Where Every Model Struggles
Drilling into the six use cases reveals glaring weaknesses:
- Screening & Discovery: Every system scores below 35%. AI agents struggle to match nuanced investment criteria across hundreds of companies and weigh qualitative factors like "tariff exposure" against quantitative filters.
- Sector, Industry & Macro: Another universal weakness — causal reasoning about macroeconomic trends and their second-order effects on specific sectors remains an open problem.
- Financial Data Extraction: The strongest category across the board, with leading systems reaching 60–70%. Models can pull numbers from filings reliably, but still miss context.
- Forward-Looking Information: Rubrics requiring predictions or scenario analysis score poorly across every system tested, confirming that "judgment" remains AI's hardest gap to close.
The takeaway is clear: finance AI is useful for data grunt work but nowhere near ready to replace human judgment for strategic decisions. The FrontierFinance benchmark proves that even the most advanced models — Claude Fable 5 included — have a long way to go before they can be trusted with the full range of an investor's workflow.
If you work in quantitative finance, AI product management, or fintech, FrontierFinance is now available as an open dataset on Hugging Face with public grading code on GitHub. The benchmark website includes interactive charts where you can compare any system combination by cost, latency, and rubric qualification rate across use cases.
Comments