STAR Framework Reveals State-Dependent Safety Failures in Multi-Turn LLM Interactions
A new study from arXiv (paper 2603.15684) introduces STAR, a state-oriented diagnostic framework that reveals how safety failures in large language models emerge from contextual state evolution across multiple conversational turns, rather than from isolated prompts. The researchers demonstrate that current safety alignment evaluations, which typically test single queries, fail to capture vulnerabilities that arise as dialogue history accumulates. STAR treats conversation history as a state transition operator, enabling controlled analysis of safety boundary traversal under autoregressive conditioning. The framework provides a principled probe for understanding how aligned models behave along interaction trajectories, without optimizing attack strength. The findings highlight a fundamental gap in existing safety alignment methodologies for multi-turn interactions.
Key facts
- arXiv paper 2603.15684 introduces STAR framework
- STAR treats dialogue history as a state transition operator
- Safety failures arise from contextual state evolution in multi-turn interactions
- Current safety alignment evaluations use isolated queries
- STAR does not optimize attack strength but probes safety boundary traversal
- Study focuses on frontier language models
- Multi-turn jailbreaks are empirically effective but poorly understood
- STAR enables controlled analysis of safety behavior along interaction trajectories
Entities
Institutions
- arXiv