HANDBOOK.md: Benchmark for Long-Context Agentic Instruction Following
Researchers have unveiled HANDBOOK.md, a benchmark featuring 65 agentic tasks designed to evaluate whether language-model agents follow long-term, binding policy documents during extended tool usage. Each task immerses an agent in a simulated corporate setting, complete with mock email, chat, calendar, issue-tracking, and commerce services accessible through the Model Context Protocol. Agents are tasked with performing routine professional duties according to expert-crafted standard operating procedures that vary from 20 to 124 pages. The tasks cover five areas, including finance and medical. This benchmark specifically fills a void in current assessments, which generally focus on task completion rather than ongoing adherence to established guidelines.
Key facts
- HANDBOOK.md consists of 65 agentic tasks
- Tasks are modeled on how enterprise employees follow company handbooks
- Each task uses a self-contained company environment with mock services
- Services include email, chat, calendar, issue-tracking, and commerce
- Services are exposed over the Model Context Protocol
- Standard operating procedures are 20 to 124 pages long
- Tasks span five domains including finance and medical
- Benchmark tests whether long policy documents constrain agent behavior
Entities
—