Harness-IF: New Benchmark Evaluates Instruction Following in Coding Agents
Researchers have introduced Harness-IF, a novel benchmark designed to evaluate instruction-following capabilities in coding agents by scoring individual operational rules based on execution evidence. The benchmark comprises 60 realistic multi-turn coding items derived from a library of 642 rules, with 256 rules receiving verdicts across five configurable instruction surfaces that deployed agents read. To distinguish genuine compliance from coincidental behavior, the team developed Against-Prior Accuracy (AP-Acc), which focuses on rules that oppose unprompted defaults, determined by re-running tasks without the rule across nine probe builds. Testing on 12 frontier models revealed accuracy ranging from 72.1% to 85.9%, while AP-Acc scores were lower, between 66.1% and 78.6%. Notably, every model performed worse on against-prior rules, with a mean drop of 5.81 points (range 3.6 to 7.4), and this trend persisted across common configurations. The study highlights that existing benchmarks often conflate rule-following with default behavior, as rules are typically concentrated in user turns and benchmarks emphasize final task success. Harness-IF addresses this gap by isolating rule adherence, providing a more precise measure of instruction following. The findings underscore the challenge of ensuring coding agents genuinely follow instructions, particularly when those instructions conflict with default behaviors. The research is available on arXiv under the identifier 2608.11727.
Key facts
- Harness-IF is a new benchmark for evaluating instruction following in coding agents
- It scores operational rules one at a time from execution evidence
- The benchmark includes 60 realistic multi-turn coding items from a 642-rule library
- 256 rules receive verdicts across five configurable instruction surfaces
- Against-Prior Accuracy (AP-Acc) separates compliance from coincidence
- AP-Acc uses nine probe builds with rules withheld to observe defaults
- 12 frontier models were tested, with accuracy ranging from 72.1% to 85.9%
- AP-Acc scores ranged from 66.1% to 78.6%
- All models performed worse on against-prior rules by 3.6 to 7.4 points (mean 5.81)
- The research is available on arXiv under identifier 2608.11727
Entities
Institutions
- arXiv