Leak-Resistant Unlearning: Benchmarking Multi-Hop Reasoning Consistency and Recovery Robustness in LLMs
A novel benchmark named Leak-Resistant Unlearning has been established to assess the effectiveness of machine unlearning techniques in large language models (LLMs). This benchmark tackles two significant issues: the risk of knowledge leakage via various multi-hop reasoning paths and the susceptibility of unlearning to recovery attacks. Previous benchmarks have mainly focused on single-hop queries and a limited array of multi-hop questions, which do not adequately evaluate the thoroughness of knowledge removal. The new benchmark evaluates models through a range of reasoning paths and recovery attacks, including lightweight adaptation after unlearning. Tests were performed on three models, six unlearning techniques, and two meticulously selected datasets. Findings reveal that current unlearning methods are at risk from these challenges. This benchmark aspires to offer a more thorough assessment of unlearning efficacy, ensuring sensitive information is genuinely erased and irretrievable. The paper can be found on arXiv under ID 2608.04519.
Key facts
- New benchmark called Leak-Resistant Unlearning introduced
- Addresses multi-hop reasoning consistency and recovery robustness
- Current benchmarks include mainly single-hop questions
- Two challenges: knowledge leakage via multi-hop paths and fragile unlearning
- Experiments on 3 models, 6 unlearning methods, 2 datasets
- Results show existing methods are vulnerable
- Benchmark evaluates recovery attacks like lightweight post-unlearning adaptation
- Paper available on arXiv:2608.04519
Entities
Institutions
- arXiv