ARTFEED — Contemporary Art Intelligence

LLM Safety Benchmarks Sensitive to Surface Form Variations

ai-technology · 2026-08-06

A recent preprint available on arXiv (2608.02665) questions the effectiveness of safety benchmarks for large language models (LLMs), positing that reliance on a single standard prompt fails to capture variations in surface forms. The authors assert that while benchmark scores serve as measurement tools, the majority of test items exist in just one standard format. They explore safety in a high-stakes context lacking a definitive gold standard, employing refusal-free techniques such as machine back-translation. Human-anchored evaluations, conducted by a judge named Claude, yielded a kappa score of 0.86 regarding unsafe compliance. Analyzing 370 seeds, 5 surface forms, and 5 models, the research indicates that no single surface form reliably predicts model safety, implying current evaluation methods may inaccurately assess LLM safety. The paper is classified as a cross-post.

Key facts

  • Preprint on arXiv (2608.02665) examines LLM safety benchmark sensitivity to surface-form variations.
  • Authors question whether single canonical prompts faithfully estimate model behavior.
  • Focus on safety, a high-stakes setting with no gold label.
  • Reformulations pre-authored using machine back-translation and Matrix-Language-Frame code-switch generator.
  • All responses scored with Claude as judge (kappa = 0.86 vs. human on unsafe compliance).
  • Judge cross-checked by GPT-4o and stable across languages.
  • Study uses 370 seeds, 5 surface forms, and 5 models.
  • No single surface form consistently predicts model safety behavior.

Entities

Institutions

  • arXiv

Sources