ARB Benchmark Exposes AI-Text Detector Weakness Against Rewritten Content
A new dataset called the Authorship-Rewriting Benchmark (ARB) has been launched by researchers to assess AI text detectors in realistic scenarios where human-written material is altered by large language models (LLMs). This benchmark fills a significant void in current evaluation techniques, which often pit human text against unmodified LLM outputs, neglecting the effects of rewriting. Comprising 1,800 human texts—600 each from XSum, WritingPrompts, and OpenWebText—ARB employs four open-weight generators: Llama-3.2-3B, Qwen2.5-7B, Mistral-7B, and Gemma-2-9B. The study assessed five detectors—FastDetectGPT, Binoculars-falcon-7b, RADAR, BERT-Defense, and RoBERTa-Defense—at a stringent 1% false-positive rate, revealing that traditional benchmarks do not accurately predict detector performance when human content is rewritten by LLMs. This research has significant implications for academic integrity and content moderation. The ARB dataset is publicly accessible for future detector advancements.
Key facts
- ARB benchmark introduced for AI-text detector evaluation
- Built from 1,800 human source texts (600 each from XSum, WritingPrompts, OpenWebText)
- Uses four open-weight generators: Llama-3.2-3B, Qwen2.5-7B, Mistral-7B, Gemma-2-9B
- Each source yields four matched variants: HUMAN, Free-LLM, H2L, LLM2L
- Evaluated five detectors: FastDetectGPT, Binoculars-falcon-7b, RADAR, BERT-Defense, RoBERTa-Defense
- Detector performance on conventional benchmarks does not predict behavior on rewritten human text
- Strict 1% false-positive rate used in evaluation
- ARB dataset is publicly available
Entities
Institutions
- arXiv
- XSum
- WritingPrompts
- OpenWebText
- Llama
- Qwen
- Mistral
- Gemma
- FastDetectGPT
- Binoculars
- RADAR
- BERT-Defense
- RoBERTa-Defense