ARTFEED — Contemporary Art Intelligence

TRACE Bench: A Task-Driven Checklist Framework for Roleplay Evaluation

ai-technology · 2026-08-13

Researchers have introduced TRACE Bench, a new framework designed for evaluating roleplay models more thoroughly than just using a single score. It takes each role profile and turns it into a detailed offline checklist. Then, by using a User Agent, it interacts with the model, updating the checklist based on the model's responses. This approach allows researchers to trace scores back to specific items and dialogue exchanges instead of relying on a general impression. In their cross-validation, they analyzed M2 free-dialogue transcripts from the MiniMax Role-play Benchmark and found only 73.74% alignment with important role-profile points, whereas TRACE Bench achieved an impressive 99.91% with fewer interactions. The findings are documented in a paper on arXiv (2608.11236), showing the framework's reliability for evaluating roleplay AI systems.

Key facts

  • TRACE Bench is a task-driven agentic checklist evaluation framework for roleplay models.
  • It decomposes each role profile into a fixed checklist offline.
  • A User Agent converses naturally with the target roleplay model while privately updating checklist states.
  • Scores trace back to checklist items and supporting dialogue turns.
  • Cross-validation used M2 free-dialogue transcripts from the MiniMax Role-play Benchmark.
  • Released free-chat transcripts covered only 73.74% of key role-profile points.
  • TRACE Bench reached 99.91% coverage in fewer turns.
  • Robustness experiments showed stable rankings under repeated runs.
  • The paper is available on arXiv with identifier 2608.11236.

Entities

Institutions

  • arXiv
  • MiniMax

Sources