ARTFEED — Contemporary Art Intelligence

MameLoshnLM: First Open-Source Yiddish Language Model and Benchmark

ai-technology · 2026-08-07

The first open-source language model tailored for Yiddish, named MameLoshnLM, has been unveiled by researchers, featuring 8 billion parameters. This initiative aims to tackle the lack of dependable digital resources for Yiddish, which has impeded advancements in language modeling, despite the language's rich literary heritage. To bolster this model, the team developed Oytser, a high-quality pretraining corpus that merges modern web-native content with literary works, alongside Kashes, a benchmark that encompasses tasks such as translation, linguistic analysis, information extraction, and language comprehension. MameLoshnLM was refined through continued pretraining on Llama 3.1 8B using these datasets. Evaluations show it surpasses comparable open baselines. The findings are published in a paper on arXiv (arXiv:2608.05850).

Key facts

  • MameLoshnLM is the first open-source 8B-parameter language model for Yiddish.
  • Oytser is a new Yiddish pretraining corpus combining web and literary sources.
  • Kashes is a multi-task benchmark for Yiddish language tasks.
  • The model is built by continuing pretraining on Llama 3.1 8B.
  • MameLoshnLM outperforms open baselines of similar scale.
  • The paper is available on arXiv with ID 2608.05850.
  • Existing multilingual corpora often contain noisy or misclassified Yiddish text.
  • The work aims to fill gaps in Yiddish language modeling resources.

Entities

Institutions

  • arXiv

Sources