Fair-ASR: A New Protocol for Comparable Black-Box Jailbreak Evaluation
On August 26, 2028, a research paper was released on arXiv (arXiv:2608.17360v1) that challenges the conventional attack success rate (ASR) metric used to assess jailbreak attacks on large language models (LLMs), claiming it fails to account for resource usage. Current evaluations that consider compute, such as FLOPs, are criticized as insufficient for black-box models. The authors propose a new evaluation framework called Fair-ASR, which restricts all approaches to a common budget of target calls (B), enabling equitable comparisons without requiring internal model insights. Upon re-assessing eleven attack strategies, the researchers observed that rankings varied significantly based on budget constraints. They found that simple stochastic perturbations and manually crafted templates were competitive, indicating that intricate attacks may not always be necessary. The findings highlight the importance of evaluation techniques that align with real-world limitations.
Key facts
- Paper published on arXiv under identifier 2608.17360v1
- Introduces Fair-ASR, an evaluation protocol for black-box jailbreak attacks
- Uses shared target-call budgets B as a comparison axis
- Target calls are directly observable and method-agnostic
- Tracks attacker calls separately for efficiency analysis
- Re-evaluates 11 representative attack methods
- Attack rankings change substantially across different target-call budgets
- Simple stochastic perturbations and hand-crafted templates remain highly competitive
Entities
Institutions
- arXiv