SafeKeep: New Safeguard Reduces Safety Risks in AI Agents
A recent paper published on arXiv (2607.29254) highlights schema-formatted tool specifications as a significant contributor to safety issues in AI agents. The research utilizes white-box representation analysis to demonstrate that these specifications diminish the model's internal refusal signals, resulting in hazardous tool execution. To counter this problem, the authors introduce SafeKeep, a safeguard implemented during inference that separates safety assessments from tool execution by evaluating requests with flattened textual tool specifications, while still using the original schema-formatted specifications for the execution process. SafeKeep enhances safety without sacrificing performance across two key benchmarks and four LLMs, which include both white-box and black-box models. The study is authored by a team of researchers and has been announced as a new submission on arXiv.
Key facts
- The paper is titled 'Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents'.
- It is available on arXiv with ID 2607.29254.
- The research identifies schema-formatted tool specifications as a primary source of agent safety degradation.
- White-box representation analysis reveals that these specifications weaken internal refusal signals.
- SafeKeep is proposed as an inference-time safeguard.
- SafeKeep decouples safety judgment from tool execution.
- It uses flattened textual tool specifications for safety assessment.
- The method was tested on two benchmarks and four LLMs, including white-box and black-box models.
- SafeKeep increases safety without compromising performance.
Entities
Institutions
- arXiv