Coding Agents Under Security Policy: Performance Losses and Cost Inflation
An arXiv preprint (2608.02670) assesses 12 coding agents using Terminal-Bench 2.1, focusing on nested security policy levels inspired by typical enterprise restrictions. The research reveals that enhancing security comes with costs, which are inconsistent: under the most stringent policy, success losses can hit 18.3 points, with cost inflation soaring to 167.3%. There is a conflict between two metrics; the model that maintains success is the one that suffers the greatest loss in efficiency, indicating that model selection is influenced by policy. This study fills a gap in current benchmarks that mainly test agents in lenient environments, leaving the impact of enforced policies unclear. The results are significant for implementing coding agents in security-focused organizations.
Key facts
- 12 coding agents evaluated on Terminal-Bench 2.1
- Nested security policy levels derived from real-world enterprise restrictions
- Strictest policy: success losses reach 18.3 points
- Cost inflation reaches 167.3% under strictest policy
- Model that best preserves success loses most efficiency
- Model choice is policy-dependent
- Existing benchmarks evaluate agents in permissive sandboxes
- Study decomposes failures induced by hardening
Entities
—