GPTKB 2.0: Million-Scale Disambiguated Knowledge Base from LLMs
Researchers have introduced GPTKB 2.0, a methodology for constructing disambiguated knowledge bases directly from large language models (LLMs). The approach addresses a core challenge in Automated Knowledge Base Construction (AKBC): LLMs lack native entity representations, leading to duplicate entries and conflations. GPTKB 2.0 incorporates on-the-fly disambiguation of entities, relations, and classes, and is designed to balance scalability with disambiguation accuracy. The system was executed at scale, producing a knowledge base with over 1 million disambiguated entities and 38.4 million triples, marking the first million-scale LLM-native KB with explicit internal canonicalization. The work is detailed in a paper on arXiv (arXiv:2608.03729v2), categorized as a cross-type announcement. The paper analyzes central design decisions and trade-offs between accuracy, scale, and cost.
Key facts
- GPTKB 2.0 is a methodology for constructing disambiguated knowledge bases from LLMs.
- It addresses duplicate entries and conflations in LLM-generated knowledge bases.
- The system performs on-the-fly disambiguation of entities, relations, and classes.
- It is designed for scalability and disambiguation accuracy.
- The materialized KB contains over 1 million disambiguated entities and 38.4 million triples.
- This is the first million-scale LLM-native KB with explicit internal canonicalization.
- The paper is available on arXiv with identifier 2608.03729v2.
- The announcement type is 'cross'.
Entities
Institutions
- arXiv