ContextWeave: New Benchmark Tests AI Memory in Real-World Office Workflows
A team of researchers has unveiled ContextWeave, a longitudinal benchmark aimed at assessing whether the recall of experiences enhances the performance of agents in realistic office workflows. This benchmark reconstructs privacy-protected, multi-month workflows from 14 participants into 1,005 executable tasks, which include 568 core evaluation tasks, complete with instructions, containerized environments, trajectories, and task-specific rubrics. It evaluates workspace quality and alignment with individual participant preferences, along with diagnostics for relevance, continuity, solvability, and resilience against misleading recall. Utilizing a fixed model across six memory components, the optimal configuration boosts the Workspace Score from 68.08 to 78.20 and the Preference Score from 41.50 to 70.60. This research fills a gap in current evaluations that typically limit memory to retrieval or question answering, highlighting its importance in long-term, stateful workflows. The benchmark can be accessed on arXiv under the identifier 2608.04830.
Key facts
- ContextWeave is a longitudinal benchmark for evaluating memory in language agents.
- It reconstructs multi-month workflows of 14 participants into 1,005 executable tasks.
- The benchmark includes 568 core evaluation tasks.
- It measures workspace quality and alignment with participant-specific preferences.
- The strongest configuration raises Workspace Score from 68.08 to 78.20.
- Preference Score improves from 41.50 to 70.60 with the strongest configuration.
- The benchmark includes diagnostics of relevance, continuity, solvability, and robustness to misleading recall.
- The work is announced on arXiv with identifier 2608.04830.
Entities
—