DharmaOCR Outperforms Newer Models on Brazilian Portuguese OCR
DharmaOCR, a specialized optical character recognition model for Brazilian Portuguese, outperforms newer generalist models Mistral OCR4 and Unlimited-OCR on Portuguese benchmarks. The model achieves a 0.925 extraction quality score versus 0.798 and 0.7587 respectively. Its advantage stems from a two-stage training pipeline: supervised fine-tuning on Portuguese documents followed by Direct Preference Optimization to ensure stability. The model correctly handles culturally specific references like Chico Buarque, which competitors misread. The structural principle is that specialization concentrates finite resources on a single domain, producing superior results even against larger, newer architectures. The benchmark results confirm that domain-specific training remains advantageous as AI advances.
Key facts
- DharmaOCR scored 0.925 on Portuguese benchmark.
- Mistral OCR4 scored 0.798.
- Unlimited-OCR scored 0.7587.
- Training used supervised fine-tuning and Direct Preference Optimization.
- Mistral OCR4 misread 'Chico Buarque' as 'Chico Barque'.
- Unlimited-OCR misread it as 'chico bique'.
- DharmaOCR handles small-font documents without degeneration.
- The paper was published three months ago on arXiv.
Entities
Artists
Institutions
- Dharma AI
- Mistral
- Unlimited-OCR
- Hugging Face
Locations
- Brazil