ARTFEED — Contemporary Art Intelligence

MiLMMT-46-v1.0: Reference-Free Post-Training Boosts Multilingual Translation

ai-technology · 2026-08-13

Researchers have launched MiLMMT-46-v1.0, a large language model designed specifically for multilingual machine translation. This new model outperforms established systems like Google Translate, Gemini 3 Pro, and GPT-5 based on reference-free quality metrics. It builds on the previous version, MiLMMT-46-v0.1, by employing Group Relative Policy Optimization (GRPO), which combines rewards from two quality estimation models, filtered by language. The model’s final version integrates SFT and RL checkpoints through linear interpolation. It shows enhanced translation quality in 46 languages compared to earlier models and exceeds recent benchmarks like Seed-X, HY-MT2, and TranslateGemma. Additionally, the research investigates on-policy distillation (OPD), which matches but doesn’t surpass the quality from RL. This work is available in arXiv paper 2608.10812.

Key facts

  • MiLMMT-46-v1.0 is an open large language model for multilingual translation.
  • It uses reference-free post-training with GRPO.
  • Reward averages two reference-free quality estimation models and is gated by language identification.
  • Model is created by linearly interpolating SFT and RL checkpoints.
  • Improves translation quality across 46 languages over SFT counterparts.
  • Outperforms Seed-X, HY-MT2, and TranslateGemma.
  • Achieves leading reference-free scores against Google Translate, Gemini 3 Pro, and GPT-5.
  • On-policy distillation reaches but does not surpass RL quality frontier.

Entities

Institutions

  • arXiv
  • Google Translate
  • Gemini 3 Pro
  • GPT-5
  • Seed-X
  • HY-MT2
  • TranslateGemma

Sources