TQLite: Multi-LLM Jury Distillation for Efficient MQM Translation Quality Evaluation
A recent preprint on arXiv (2608.02975) presents TQLite, a framework for distillation that allows small language models (SLMs) to nearly match the performance of large reasoning models (LRMs) in MQM-based translation quality (TQ) assessments. The research evaluates SLMs, LLMs, and LRMs across different TQ evaluation configurations, outlining optimal practices. TQLite utilizes a multi-LRM jury to produce high-quality synthetic training data through effective data curation strategies, tackling the scalability issues associated with the use of resource-intensive LLMs and LRMs. This work is crucial for enhancing the efficiency and accessibility of advanced translation quality evaluations.
Key facts
- arXiv:2608.02975v1
- Announce Type: cross
- Large language models (LLMs) show impressive performance in MQM-based translation quality evaluation
- Large reasoning models (LRMs) promise further improvements
- LLMs and LRMs are computationally expensive to deploy at scale
- Small language models (SLMs) are more efficient but struggle with complex reasoning
- Extensive empirical study benchmarks SLMs, LLMs, and LRMs across TQ evaluation setups
- TQLite is a novel distillation framework
- TQLite enables SLMs to approach the MQM evaluation performance of the best LRM-based evaluators
- Approach uses a multi-LRM jury to generate high-quality synthetic training data
- Practical data curation techniques are used
Entities
Institutions
- arXiv