GAUGE Benchmark Evaluates AI Financial Models Against Analyst Practice
A new benchmark named GAUGE has been developed by researchers to assess financial valuation models created by agents, relying on actual analyst practices instead of a single expert's evaluation. The investigation covered 108 directed pairs from 65 companies, revealing a median score of 0.33 for single-reference evaluations, with 92.6% of the scores below 0.70. Additionally, no pairs from the same vintage agreed on implied prices within a 10% margin. This suggests that point-tolerance grading reflects the existing discrepancies among professionals. GAUGE incorporates 1,001 analyst workbooks categorized by vendors and utilizes a 196-task evaluation framework featuring a three-layer observed-practice structure. This research was published on arXiv under ID 2607.24889v1.
Key facts
- GAUGE stands for Grading Agent-Built Financial Models Without a Golden Answer
- Study used 108 directed pairs covering 65 companies
- Median single-reference score was 0.33
- 92.6% of scores were below 0.70
- No same-vintage pair agreed on implied price within 10%
- Benchmark uses 1,001 vendor-classified analyst workbooks
- Evaluation set includes 196 tasks
- Published on arXiv with ID 2607.24889v1
Entities
Institutions
- arXiv