ARTFEED — Contemporary Art Intelligence

SRE-Bench: First Contamination-Free Reverse Engineering Benchmark for AI Agents

ai-technology · 2026-08-13

A new benchmark, SRE-Bench, has been introduced to evaluate AI agents' reverse engineering (RE) capabilities on binary software. The benchmark addresses a critical gap: while AI agents excel at analyzing source code, much consequential software (malware, firmware, proprietary apps) is only available as binaries, requiring RE to recover program semantics. Existing benchmarks fail to prevent contamination (instances seen in LLM training data) and lack real-world scale and anti-analysis protections. SRE-Bench, built from scratch by RE experts with over 5,000 hours of effort, is the first realistic, contamination-free RE benchmark. It ensures instances are unseen as source code in LLM training data, forcing genuine analysis rather than recognition. The benchmark matches the scale and anti-analysis protections of real software, providing a rigorous test for agentic cybersecurity. The paper is available on arXiv (2608.11469).

Key facts

  • SRE-Bench is the first realistic, contamination-free reverse engineering benchmark for AI agents.
  • It was built entirely from scratch by RE experts with over 5,000 hours of work.
  • The benchmark addresses the challenge of evaluating agentic RE on binaries, which are common in malware, firmware, and proprietary software.
  • Existing benchmarks do not jointly satisfy requirements of being unseen in LLM training data and matching real software scale/protections.
  • SRE-Bench prevents models from taking shortcuts by recognizing source code from training data.
  • The benchmark is designed to test AI agents' ability to recover program semantics from binaries.
  • The paper is announced on arXiv with ID 2608.11469.
  • The work is relevant to cybersecurity, where reverse engineering is critical for analyzing malicious software.

Entities

Institutions

  • arXiv

Sources