Why the Most Powerful AI Is Not Always Best for Translation
Short answer: the newest frontier models are tuned to win at step-by-step reasoning (maths, code, logic). Translation rewards something else: fluent, natural writing in the target language. Published research finds that switching on explicit reasoning consistently lowers machine translation quality, so more reasoning is not more translation quality.
What the research shows
- Rajaee et al. (2026) tested several reasoning models on the WMT24++ benchmark and found that enabling explicit reasoning consistently degrades translation quality across languages and models.
- Li et al. (2025) found that chain-of-thought reasoning lowers instruction-following accuracy across 15 models.
- Liu et al. (2026) found that mainstream reasoning models gain little on open-ended writing compared with maths.
- Zhang et al. (ACL 2026 Findings) found that strong comprehension does not produce human-level creativity in literary translation.
How TranslationAI chooses models
We pick models for writing quality per cost, and run translation at a low reasoning setting so the model spends its effort on the sentence, not on deliberation.
Further reading: choose a model and framework, quality vs speed vs cost.