ARTFEED — Contemporary Art Intelligence

GPTKB 2.0: A Disambiguated LLM-Derived Knowledge Base with Web Demo

ai-technology · 2026-08-06

A new large-scale knowledge base, GPTKB 2.0, has been introduced by researchers, featuring a web demo for users to browse, query, and review its information. This knowledge base encompasses 38.4 million triples linked to 1.6 million canonical entities, including 207,600 consolidated relations and 66,000 consolidated classes. In contrast to earlier LLM-based knowledge bases that typically recognized entities through surface strings, GPTKB 2.0 utilizes context-aware disambiguation during its recursive construction, effectively distinguishing homonyms and combining synonymous mentions. The demo allows users to explore entities, navigate links within the KB, and verify the origins of specific facts, including surface forms and disambiguation choices. The interface accommodates structured SPARQL queries and natural-language questions. This project is available on arXiv under the identifier 2608.06992, marking a notable advancement in enhancing the transparency and reliability of LLM-derived knowledge for research and practical use.

Key facts

  • GPTKB 2.0 contains 38.4M triples over 1.6M canonical entities.
  • It includes 207.6K consolidated relations and 66K consolidated classes.
  • The knowledge base uses context-guided disambiguation to separate homonyms and merge synonyms.
  • A web demo allows browsing, querying, and auditing of the KB.
  • Users can audit provenance of facts, including surface forms and disambiguation decisions.
  • The interface supports SPARQL queries and natural-language questions.
  • Entity linking from user-provided text to canonical entries is supported.
  • The project is available on arXiv (2608.06992).

Entities

Institutions

  • arXiv

Sources