Magnet: Detecting Cross-Session AI Misuse via Capability Accumulation
A new arXiv preprint (2608.02518) introduces Magnet, a framework designed to detect cross-session AI misuse, a threat model where attackers decompose harmful goals into innocuous subtasks executed across isolated agentic sessions. The authors argue that while most AI abuse detection focuses on single-turn or multi-turn (single-session) attacks, this approach overlooks the asymmetry where the AI is stateless between conversations but the attacker is not, enabling evasion. The paper demonstrates cross-session goal decomposition as an effective evasion technique, potentially eliciting more harmful capability than equivalent single-session attacks. The work addresses a critical gap in AI safety, proposing a detection method based on capability accumulation across sessions. The preprint was announced on arXiv with the identifier 2608.02518.
Key facts
- The paper is titled 'Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation'.
- It is an arXiv preprint with identifier 2608.02518.
- The research focuses on cross-session AI misuse, a threat model not covered by existing detection frameworks.
- Attackers can decompose harmful goals into innocuous units and execute them in isolated agentic sessions.
- The agent is stateless between conversations, but the attacker is not, creating an asymmetry.
- The paper demonstrates cross-session goal decomposition as an evasion technique.
- It may elicit more harmful capability than equivalent single-session or multi-turn attacks.
- The work proposes a detection method based on capability accumulation.
Entities
Institutions
- arXiv