ARTFEED — Contemporary Art Intelligence

FinProBench: New Benchmark for Financial AI Agents with Role-Grounded Rubrics

ai-technology · 2026-08-06

FinProBench has been launched by researchers as a benchmark to assess financial AI agents, utilizing criteria based on professional outputs. Accompanying this benchmark is the Role-Grounded Rubric Construction (RGRC), a reusable pipeline that derives evaluation rubrics from the actual deliverables of professionals in defined roles. RGRC is structured into four phases: Deliverable Collection, Competency Extraction, Rubric Synthesis, and Validation. The rubrics created through RGRC encapsulate implicit standards, differentiate quality levels, and are transferable across tasks within a role. In their study, the researchers categorized 57 occupations by deliverable type into 30 conventional roles with rich prior knowledge and 27 specialized roles with sparse prior knowledge. The Prompt-only method closely approaches RGRC for conventional roles (89.2% compared to 90.7%), yet RGRC significantly excels in role-specialized categories. The research paper can be found on arXiv with the identifier 2608.04077.

Key facts

  • FinProBench is a new benchmark for professional financial tasks.
  • RGRC (Role-Grounded Rubric Construction) is a reusable pipeline for deriving rubrics from practitioner deliverables.
  • RGRC comprises four stages: Deliverable Collection, Competency Extraction, Rubric Synthesis, and Validation.
  • The rubrics capture tacit standards, distinguish quality levels, and transfer across tasks within a role.
  • 57 occupations were classified into 30 conventional roles and 27 role-specialized roles.
  • Prompt-only nearly matches RGRC for conventional roles (89.2% vs. 90.7%).
  • RGRC substantially outperforms Prompt-only for role-specialized roles.
  • The paper is available on arXiv with identifier 2608.04077.

Entities

Institutions

  • arXiv

Sources