Trait-Invariant Safety Tuning Stabilizes LLM Safety Across System Prompt Traits
A recent preprint on arXiv (2608.11705) presents Trait-Invariant Safety Tuning (TIST), a technique designed to enhance the safety performance of large language models (LLMs) across various traits specified in system prompts. The researchers illustrate that identical user requests can result in varying safety responses based on the indicated trait (such as personality or role), a phenomenon referred to as 'trait-induced safety variation.' To measure this, they introduce two refusal-based metrics: Trait-Induced Deviation, which assesses dataset-level divergence from a no-trait baseline, and Trait-Induced Flip Rate, which evaluates the consistency of safety decisions across traits for the same request. Their analysis reveals that traits disrupt the model's safety representations within a low-dimensional subspace. TIST seeks to ensure uniform safety behavior, independent of the trait assigned. This paper has been announced as a new submission on arXiv.
Key facts
- The paper is titled 'Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning'.
- It is available on arXiv with identifier 2608.11705.
- The failure mode is called 'trait-induced safety variation'.
- Two metrics are introduced: Trait-Induced Deviation and Trait-Induced Flip Rate.
- Traits perturb safety representations within a low-dimensional subspace.
- The proposed method is Trait-Invariant Safety Tuning (TIST).
- The goal is to stabilize safety behavior across traits.
- The paper is a new submission (Announce Type: new).
Entities
Institutions
- arXiv