ARTFEED — Contemporary Art Intelligence

FinPerMA: New Benchmark Tests LLM Agents' Personalized Memory in Finance

ai-technology · 2026-08-06

The introduction of FinPerMA marks a significant advancement in assessing the personalized memory abilities of large language model (LLM) agents in critical areas like financial advising. This benchmark, outlined in a paper available on arXiv (arXiv:2608.04095), fills a void in current personalized-memory assessments, which mainly focus on factual retention or depend on loosely defined model-generated paths, neglecting the adaptation of preferences based on events. FinPerMA tests personalized memory through fixed longitudinal investor paths, utilizing a generation pipeline that merges deterministic, theory-based impact rules, regulated LLM narration, and automated quality checks. A Post-Shock checkpoint determines if an agent has incorporated a significant event into its enduring user model. The benchmark includes 2,994 questions from 276 personas. Testing seven advanced LLMs with various memory setups reveals that performance remains far from optimal, suggesting ample opportunity for enhancement. This research is crucial for the evolution of AI assistants in finance and other tailored fields, emphasizing the necessity for stronger memory and adaptation strategies.

Key facts

  • FinPerMA is an event-grounded benchmark for personalized memory in LLM agents.
  • It evaluates against frozen longitudinal investor trajectories.
  • The generation pipeline uses deterministic, theory-informed impact rules, controlled LLM narration, and automated quality screening.
  • A Post-Shock checkpoint isolates integration of material events into persistent user models.
  • The benchmark includes 2,994 questions from 276 personas.
  • Seven frontier LLMs and up to seven memory configurations were tested.
  • Results show performance is far from saturated.
  • The paper is available on arXiv with ID 2608.04095.

Entities

Institutions

  • arXiv

Sources