ARTFEED — Contemporary Art Intelligence

HANDBOOK.md: Benchmark for Long-Context Agentic Instruction Following

ai-technology · 2026-07-29

Researchers have unveiled HANDBOOK.md, a benchmark featuring 65 agentic tasks designed to evaluate whether language-model agents follow long-term, binding policy documents during extended tool usage. Each task immerses an agent in a simulated corporate setting, complete with mock email, chat, calendar, issue-tracking, and commerce services accessible through the Model Context Protocol. Agents are tasked with performing routine professional duties according to expert-crafted standard operating procedures that vary from 20 to 124 pages. The tasks cover five areas, including finance and medical. This benchmark specifically fills a void in current assessments, which generally focus on task completion rather than ongoing adherence to established guidelines.

Key facts

  • HANDBOOK.md consists of 65 agentic tasks
  • Tasks are modeled on how enterprise employees follow company handbooks
  • Each task uses a self-contained company environment with mock services
  • Services include email, chat, calendar, issue-tracking, and commerce
  • Services are exposed over the Model Context Protocol
  • Standard operating procedures are 20 to 124 pages long
  • Tasks span five domains including finance and medical
  • Benchmark tests whether long policy documents constrain agent behavior

Entities

Sources