TongGuOCR: AI Framework for Chinese Historical Document Transcription
A new OCR framework, TongGuOCR, has been proposed to improve the transcription of Chinese historical documents. The framework, detailed in a paper on arXiv (2608.07917), addresses challenges such as complex layouts, rare characters, and nontrivial reading orders. It consists of a Layout-Aware Preprocessing module that constructs locally coherent recognition blocks to preserve context, and a Token-Augmented Recognition module that expands the vocabulary at the character level to give rare glyphs direct token representations, shortening decoding paths. The framework also augments transcription at the line level. The paper was announced as a new submission on arXiv.
Key facts
- TongGuOCR is a layout-aware and token-augmented OCR framework for Chinese historical documents.
- The framework is described in arXiv paper 2608.07917, announced as a new submission.
- It addresses challenges in OCR for historical documents: complex layouts, rare characters, and nontrivial reading orders.
- The Layout-Aware Preprocessing module constructs and refines locally coherent recognition blocks.
- The Token-Augmented Recognition module expands vocabulary at the character level for rare glyphs.
- The framework shortens decoding paths for rare characters by giving them one-token representations.
- Line-to-line transitions are also augmented in the recognition module.
- The goal is to enable full-text retrieval, collation, and computational analysis of scanned historical documents.
Entities
Institutions
- arXiv