ARTFEED — Contemporary Art Intelligence

FinEvo-Bench: Longitudinal Benchmark for Self-Evolving Agents in Finance

other · 2026-08-07

A new longitudinal benchmark called FinEvo-Bench has been developed by researchers to assess self-evolving agents within professional financial workflows. This benchmark fills a void in current agent evaluations, which often analyze tasks in isolation and do not consider how experience from one task may enhance performance in others. FinEvo-Bench includes 120 tasks based on real cases, covering 20 business scenarios across six financial sectors. Each task is rooted in professional procedures provided by institutions, with eligible cases offering factual data. Every scenario features six related yet distinct cases, accompanied by a reviewed rubric for quality and compliance. The benchmark evaluates four self-evolving agent frameworks using the Qwen3.7-Max backbone across three interleaved task streams. Non-evolving controls are included to gauge improvements. This benchmark aims to deliver a more accurate assessment of agents in professional environments, emphasizing the importance of learning from previous tasks. The research paper can be found on arXiv under the identifier 2608.06144.

Key facts

  • FinEvo-Bench is a longitudinal benchmark for self-evolving agents in professional financial workflows.
  • It includes 120 real-case-grounded tasks across 20 business scenes and six financial domains.
  • Tasks are based on institution-provided professional procedures and publicly documented cases.
  • Each scene contains six related but distinct cases sharing a professional procedure and a rubric.
  • The benchmark compares four self-evolving agent scaffolds using the Qwen3.7-Max backbone.
  • Three independently shuffled, globally interleaved task streams are used for evaluation.
  • Paired non-evolving controls estimate each scaffold's improvement.
  • The paper is available on arXiv with identifier 2608.06144.

Entities

Institutions

  • arXiv

Sources