ForgetBench: A Benchmark for LLM Forgetting Under Continual Knowledge Editing
ForgetBench is a newly developed benchmark aimed at systematically analyzing the forgetting tendencies of large language models (LLMs) during ongoing knowledge updates. It fills a void in current evaluation methods that primarily emphasize single-step reasoning or static knowledge modifications, which overlook the temporal aspects of knowledge retention and loss throughout successive model revisions. This benchmark presents two interrelated evaluation approaches: concept-based QA and scenario-based QA, which differentiate between isolated factual retention and the preservation of structured relational knowledge. Utilizing a sequential editing framework, ForgetBench creates temporally organized knowledge streams and assesses model performance at various editing phases. This research is documented in arXiv:2607.26455.
Key facts
- ForgetBench is a benchmark for LLM forgetting under continual knowledge editing.
- It introduces concept-based QA and scenario-based QA evaluation paradigms.
- It uses a sequential editing framework with temporally ordered knowledge streams.
- Existing paradigms fail to capture temporal dynamics of knowledge retention.
- The work is published on arXiv with ID 2607.26455.
- The benchmark aims to understand knowledge degradation during continual model modification.
Entities
Institutions
- arXiv