DocAnnot: AI Framework Accelerates Document Annotation
Researchers have introduced DocAnnot, a framework that uses a Large Vision Language Model (LVLM), OCR, and a Spatially Informed Contextual Matching (SICM) algorithm to automate the creation of Key Information Extraction (KIE) datasets. Tested on CORD and SROIE benchmarks, it achieved F1-scores of 0.679 and 0.846. While human-annotated data remains superior, models fine-tuned solely on DocAnnot's outputs showed respectable performance, significantly reducing manual effort.
Key facts
- DocAnnot uses LVLM, OCR, and SICM algorithm
- SICM combines spatial relationships, proximity analysis, and textual matching
- Evaluated on CORD and SROIE benchmarks
- F1-scores: 0.679 on CORD, 0.846 on SROIE
- Auto-annotated data can be used for fine-tuning downstream KIE models
- Human-annotated data still outperforms auto-annotated data
- Framework aims to accelerate KIE dataset generation
- Published on arXiv with ID 2607.24745
Entities
Institutions
- arXiv