ARTFEED — Contemporary Art Intelligence

New Benchmark for Fine-Grained Attribution in Long-Horizon LLM Agent Trajectories

ai-technology · 2026-08-10

A recent research paper presents a comprehensive benchmark and annotation framework aimed at trajectory attribution for long-horizon large language model (LLM) agents. This study, accessible on arXiv (2608.06909), tackles the shortcomings in detailed attribution analysis found in current agent benchmarks, which mainly focus on behavioral results. The framework categorizes diverse agent trajectories using a standardized component schema and offers annotations for the key attribution components, including relevant attack and execution sequences. It features over 1,300 annotated trajectories from AgentDojo and the Stage and Canary environments of Agent3Sigma, encompassing task-related actions, unsafe actions, and safety refusals. Two evaluation tasks are defined: primary attribution localization and attribution-chain recovery, complete with reference implementations. This advancement is crucial for the AI research community, facilitating a more profound examination of agent decision-making and safety, which may guide future enhancements and safety protocols.

Key facts

  • The paper introduces a benchmark and annotation framework for trajectory attribution in LLM agents.
  • It is available on arXiv with identifier 2608.06909.
  • The benchmark uses a unified component schema for heterogeneous trajectories.
  • Annotations include primary attribution component, attack chains, and execution chains.
  • Trajectories are sourced from AgentDojo and Agent3Sigma (Stage and Canary settings).
  • Over 1,300 annotated trajectories are included.
  • The benchmark defines two evaluation tasks: primary attribution localization and attribution-chain recovery.
  • The work addresses limitations of existing benchmarks that focus only on behavioral outcomes.

Entities

Institutions

  • arXiv
  • AgentDojo
  • Agent3Sigma

Sources