ARTFEED — Contemporary Art Intelligence

DocAnnot: AI Framework Accelerates Document Annotation

ai-technology · 2026-07-29

Researchers have introduced DocAnnot, a framework that uses a Large Vision Language Model (LVLM), OCR, and a Spatially Informed Contextual Matching (SICM) algorithm to automate the creation of Key Information Extraction (KIE) datasets. Tested on CORD and SROIE benchmarks, it achieved F1-scores of 0.679 and 0.846. While human-annotated data remains superior, models fine-tuned solely on DocAnnot's outputs showed respectable performance, significantly reducing manual effort.

Key facts

  • DocAnnot uses LVLM, OCR, and SICM algorithm
  • SICM combines spatial relationships, proximity analysis, and textual matching
  • Evaluated on CORD and SROIE benchmarks
  • F1-scores: 0.679 on CORD, 0.846 on SROIE
  • Auto-annotated data can be used for fine-tuning downstream KIE models
  • Human-annotated data still outperforms auto-annotated data
  • Framework aims to accelerate KIE dataset generation
  • Published on arXiv with ID 2607.24745

Entities

Institutions

  • arXiv

Sources