ARTFEED — Contemporary Art Intelligence

New Safety Architecture for LLM Mental Health Support

ai-technology · 2026-07-29

A safety governance framework that is model-agnostic has been created by researchers for large language models utilized in mental health assistance. This architecture integrates contextual risk identification, verification through reasoning, and response generation guided by protocols for multi-turn dialogues. Evaluated using GPT-5-chat and Qwen3.5-27B on synthetic dialogues based on authentic mental health stories, it demonstrated excellent risk detection capabilities (specificity: 0.85, 95% CI: 0.78-0.91; sensitivity: 0.92, 95% CI: 0.88-0.95) and enhanced clinician-preferred escalation responses by 25.6-59.2 percentage points, all while maintaining rapport. The system's performance was consistent regardless of conversation length and effectively generalized across various models.

Key facts

  • Architecture is model-agnostic
  • Combines risk detection, verification, and response generation
  • Tested with GPT-5-chat and Qwen3.5-27B
  • Specificity: 0.85 (95% CI: 0.78-0.91)
  • Sensitivity: 0.92 (95% CI: 0.88-0.95)
  • Increased clinician-preferred escalation responses by 25.6-59.2 percentage points
  • Preserved rapport and connection
  • Performance stable across conversation length

Entities

Sources