REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems
A recent paper on arXiv (2608.10669) has unveiled a new standard for assessing the safety of large language model (LLM) agents. Named REDAgentBench, this executable framework facilitates autonomous red-teaming and accurate evaluation of agent systems. It generates attacks based on defined safety constraints and vulnerabilities, executing them within isolated service sandboxes while confirming harmful outcomes through service receipts and changes in final states. The benchmark encompasses 1,661 scenarios across five service surfaces. Analyzing six models and three agent harnesses reveals a macro-average attack success rate (ASR) of 65.69%. The authors contend that current assessments often oversimplify agent safety to a single ASR, merging exposure, execution, observation, and adjudication, which may obscure genuine violations. REDAgentBench seeks to enhance measurement fidelity by distinguishing these elements.
Key facts
- REDAgentBench is an executable framework for autonomous red-teaming and faithful measurement of LLM agent systems.
- It derives attacks from explicit safety constraints and associated agent-system vulnerabilities.
- Attacks are run in isolated service sandboxes.
- Harmful effects are verified from service receipts and final-state changes.
- The benchmark contains 1,661 cases across five service surfaces.
- Across six models and three agent harnesses, the macro-average ASR is 65.69%.
- The paper criticizes existing evaluations for reducing agent safety to a single ASR.
- REDAgentBench aims to separate exposure, execution, observation, and adjudication for faithful measurement.
Entities
Institutions
- arXiv