OBLIVION: New Benchmark for Unlearning Skills in AI Agents
A recent study presents OBLIVION, a benchmark and defense mechanism aimed at tackling the unlearning of operational skills in large language model (LLM) agents. Available on arXiv (ID: 2608.08264), the paper discusses the issue of agents being able to retrieve deleted skills from leftover data such as archives and memory entries. OBLIVION conceptualizes episodes as workflows from source to sink and employs Cross-Surface Coherent Erasure to reduce these residual carriers. In experiments involving 88 locked attack episodes, the no-defense scenario achieved a 1.0 attack success rate, whereas OBLIVION managed to lower this to 0.114. This research is pivotal for addressing significant security issues in AI deployment, particularly in preventing the misuse of revoked skills by agents, and is crucial for AI safety and agent-based systems.
Key facts
- Paper titled 'OBLIVION: Workflow-Level Operational Skill Unlearning for Deployed Agents'
- Posted on arXiv with ID 2608.08264
- Addresses operational skill unlearning in LLM agents
- Introduces benchmark and defense harness called OBLIVION
- Uses Cross-Surface Coherent Erasure to reduce residual carriers
- Applies frozen workflow remediation near dangerous sinks
- On 88 locked attack episodes, no-defense arm had attack success rate 1.0
- OBLIVION reduces attack success rate to 0.114
Entities
Institutions
- arXiv