ARTFEED — Contemporary Art Intelligence

FrontierFinance: New Benchmark for AI Financial Agents

ai-technology · 2026-08-13

FrontierFinance has been launched by researchers as a robust benchmark aimed at assessing AI agents within the realm of professional investment research. This benchmark, outlined in a paper on arXiv (2608.11683), seeks to overcome the shortcomings of current benchmarks that primarily concentrate on financial data extraction, an area where existing models have become saturated. FrontierFinance includes 220 expertly designed queries and 11,543 source-attributed rubrics covering six use cases throughout the entire investor workflow, making it both more comprehensive and challenging than existing public finance benchmarks. An evaluation of frontier models and agent systems, utilizing a common framework limited to publicly accessible data, demonstrated that the tool harness significantly impacts quality and efficiency. Samaya's proprietary system achieved a leading score of 56.0%, surpassing the top frontier models. This fully open benchmark aims to reflect the intricacies of genuine analyst inquiries, which tend to be open-ended and lengthy, thereby offering a more precise assessment of frontier intelligence in finance.

Key facts

  • FrontierFinance is a new benchmark for AI agents in investment research.
  • It includes 220 expert-crafted queries and 11,543 source-attributed rubrics.
  • The benchmark spans six use cases across the full investor workflow.
  • Existing benchmarks focus on financial data extraction, which models have saturated.
  • Reference-based metrics and generic LLM-as-a-judge scoring are insufficient for open-ended answers.
  • Evaluation shows the tool harness strongly shapes quality and efficiency.
  • Samaya's in-house system leads with 56.0% accuracy.
  • The benchmark is fully open and restricted to publicly available data.

Entities

Institutions

  • arXiv
  • Samaya

Sources