CREBench: Benchmarking LLMs for Cryptographic Binary Reverse Engineering
A new standard known as CREBench has been launched to assess the effectiveness of large language models (LLMs) in cryptographic binary reverse engineering (RE). This benchmark includes 432 challenges derived from 48 established cryptographic algorithms, featuring 3 scenarios of insecure crypto key usage and 3 levels of difficulty. Each challenge is framed as a Capture-the-Flag (CTF) RE task, requiring the model to interpret the cryptographic logic and retrieve the correct input. This evaluation framework aims to systematically measure LLM performance in this niche area. The research, which fills a gap in understanding LLMs' role in reverse engineering, is detailed in the arXiv paper 2604.03750, emphasizing the potential for LLMs to streamline RE processes, vital for discovering vulnerabilities in sensitive cryptographic programs.
Key facts
- CREBench is a benchmark for evaluating LLMs in cryptographic binary reverse engineering.
- It includes 432 challenges built from 48 standard cryptographic algorithms.
- The benchmark covers 3 insecure crypto key usage scenarios and 3 difficulty levels.
- Each challenge follows a Capture-the-Flag (CTF) format.
- The study is published on arXiv with ID 2604.03750.
- The announcement type is 'replace-cross'.
- The research aims to automate reverse engineering processes using LLMs.
- Reverse engineering is critical for vulnerability discovery and malware analysis.
Entities
Institutions
- arXiv