ARTFEED — Contemporary Art Intelligence

JieZi: Large-Scale Dataset and Benchmark for Ancient Chinese Character Exegesis

ai-technology · 2026-08-13

A new task known as Ancient Chinese Character Exegesis (ACCE) has been developed by researchers, focusing on vision-language question answering (VQA) to analyze ancient Chinese characters. This task is divided into four levels: identifying basic characters, analyzing glyph forms, interpreting meanings, and studying diachronic evolution. To facilitate this, the JieZi-Dataset was created, marking the first extensive, expert-verified VQA training dataset for ACCE, which includes over 500,000 question-answer pairs. The dataset was generated through a method that minimizes factual inaccuracies by leveraging expert insights. This initiative aims to fill the gap in structured datasets and benchmarks for thorough scholarly examination, as current computational methods primarily concentrate on limited tasks like character recognition. The paper can be found on arXiv with the identifier 2608.11741.

Key facts

  • Introduction of ACCE, a VQA task for ancient Chinese character exegesis
  • ACCE has four progressive levels: identification, glyph-form analysis, meaning exegesis, diachronic evolution
  • JieZi-Dataset is the first large-scale, expert-audited VQA training dataset for ACCE
  • JieZi-Dataset contains over 500,000 QA pairs
  • Dataset construction pipeline reduces factual errors by constraining generation
  • Existing computational approaches focus on subtasks like recognition and retrieval
  • Paper available on arXiv with ID 2608.11741

Entities

Institutions

  • arXiv

Sources