ARTFEED — Contemporary Art Intelligence

SteeringSafety Benchmark Evaluates Representation Steering Across Safety Perspectives

ai-technology · 2026-08-13

Researchers have introduced a new standard called SteeringSafety, aimed at evaluating steering techniques for large language models (LLMs) through nine different safety perspectives using 18 datasets. This benchmark focuses on key safety aspects such as refusal, bias, hallucination, social behavior, reasoning, epistemic integrity, and normative judgment. SteeringSafety includes modular elements for complex steering tactics, allowing for the integration of DIM, ACE, CAA, PCA, and LAT, along with new features like conditional steering. Results from models like Gemma-2-2B, Llama-3.1-8B, and Qwen-2.5-7B show that the best steering outcomes depend on the method, model, and viewpoint. Interestingly, while DIM consistently performs well, improvements in one safety area can negatively impact others, highlighting the need for careful method selection based on safety objectives.

Key facts

  • SteeringSafety is a benchmark for evaluating representation steering methods.
  • It covers nine safety perspectives across 18 datasets.
  • Safety perspectives include refusal, bias, hallucination, social behaviors, reasoning, epistemic integrity, and normative judgment.
  • SteeringSafety provides modularized building blocks for DIM, ACE, CAA, PCA, and LAT methods.
  • Recent enhancements such as conditional steering are included.
  • Results were obtained on Gemma-2-2B, Llama-3.1-8B, and Qwen-2.5-7B models.
  • DIM is consistently effective across models and perspectives.
  • All methods exhibit substantial entanglement between safety perspectives.
  • The benchmark is introduced in arXiv paper 2509.13450.
  • The paper was announced as a replace on arXiv.

Entities

Institutions

  • arXiv

Sources