ARTFEED — Contemporary Art Intelligence

ContextWeave: New Benchmark Tests AI Memory in Real-World Office Workflows

ai-technology · 2026-08-06

A team of researchers has unveiled ContextWeave, a longitudinal benchmark aimed at assessing whether the recall of experiences enhances the performance of agents in realistic office workflows. This benchmark reconstructs privacy-protected, multi-month workflows from 14 participants into 1,005 executable tasks, which include 568 core evaluation tasks, complete with instructions, containerized environments, trajectories, and task-specific rubrics. It evaluates workspace quality and alignment with individual participant preferences, along with diagnostics for relevance, continuity, solvability, and resilience against misleading recall. Utilizing a fixed model across six memory components, the optimal configuration boosts the Workspace Score from 68.08 to 78.20 and the Preference Score from 41.50 to 70.60. This research fills a gap in current evaluations that typically limit memory to retrieval or question answering, highlighting its importance in long-term, stateful workflows. The benchmark can be accessed on arXiv under the identifier 2608.04830.

Key facts

  • ContextWeave is a longitudinal benchmark for evaluating memory in language agents.
  • It reconstructs multi-month workflows of 14 participants into 1,005 executable tasks.
  • The benchmark includes 568 core evaluation tasks.
  • It measures workspace quality and alignment with participant-specific preferences.
  • The strongest configuration raises Workspace Score from 68.08 to 78.20.
  • Preference Score improves from 41.50 to 70.60 with the strongest configuration.
  • The benchmark includes diagnostics of relevance, continuity, solvability, and robustness to misleading recall.
  • The work is announced on arXiv with identifier 2608.04830.

Entities

Sources