ARTFEED — Contemporary Art Intelligence

NVIDIA Nemotron Adapted for Modern Greek: New HERA Benchmark and 0.835 nDCG@10

ai-technology · 2026-08-06

Researchers have enhanced NVIDIA's Nemotron retrieval stack by incorporating Modern Greek, filling a gap in the model and in important multilingual benchmarks. Their findings, published in a paper on arXiv (2608.05138), concentrate on retrieval-augmented generation (RAG) in fields like law, energy, finance, and healthcare. The team conducted corpus mining, implemented synthetic supervision, trained the retrieval model, adapted the reranker, and fine-tuned the reader, leading to a new benchmark called HERA. Interestingly, a parameter-free BM25 baseline outperformed several existing multilingual dense retrieval models on Greek datasets. After fine-tuning with 65,773 Greek retrieval pairs, the Nemotron 1B embedder significantly improved nDCG@10 from 0.362 to 0.835, greatly surpassing its original version.

Key facts

  • Modern Greek is absent from NVIDIA's Nemotron retrieval models and major multilingual retrieval benchmarks.
  • The adaptation targets RAG in legal, energy, financial, and medical applications.
  • The process includes corpus mining, synthetic supervision, retrieval model training, reranker adaptation, and reader fine-tuning.
  • A new benchmark called HERA was introduced.
  • BM25 baseline outperforms several off-the-shelf multilingual dense retrieval models on specialist Greek corpora.
  • Fine-tuning on 65,773 Greek retrieval pairs improved nDCG@10 from 0.362 to 0.835 for a Nemotron 1B embedder.
  • The adapted model substantially outperforms its unadapted counterpart.
  • Learned language competence transfers to general-domain Greek, but the advantage over BM25 is domain-dependent.
  • A cross-encoder reranker was also adapted.
  • The paper is available on arXiv with ID 2608.05138.

Entities

Institutions

  • NVIDIA
  • arXiv

Sources