TwT: Difficulty-Adaptive Reasoning for Multi-Domain Machine Translation
A novel approach known as Translation with Thought (TwT) has been unveiled to enhance multi-domain machine translation (MDMT) by adjusting reasoning efforts according to linguistic intricacies. Drawing inspiration from human translators, this technique alternates between intuitive and deliberate reasoning. TwT undergoes a two-phase training process: first, it involves supervised fine-tuning on difficulty-aware long chain-of-thought traces derived from DeepSeek-R1 and rephrased by GPT-4o; second, it employs reinforcement learning with a hybrid reward to refine translation quality and reasoning efficiency. Tested across 15 benchmarks in both in-domain and out-of-domain contexts, as well as on 3 familiar and 59 unfamiliar languages, TwT-7B and TwT-14B surpass larger state-of-the-art reasoning models in translation quality while minimizing token usage. The research can be found on arXiv under ID 2607.29287.
Key facts
- TwT (Translation with Thought) is a resource-rational framework for multi-domain machine translation.
- It adapts reasoning effort based on linguistic complexity, inspired by human translators.
- Training involves supervised fine-tuning on difficulty-aware long chain-of-thought traces from DeepSeek-R1 and rewritten by GPT-4o.
- Reinforcement learning with a hybrid reward optimizes translation quality and reasoning efficiency.
- Evaluated on 15 benchmarks, including in-domain and out-of-domain settings.
- Tested on 3 seen and 59 unseen languages.
- Ablations across three backbone models were performed.
- TwT-7B and TwT-14B outperform larger SOTA reasoning models in translation quality while reducing token usage.
Entities
Institutions
- arXiv
- DeepSeek
- OpenAI