North Small Translate’s Agentic Model Hits 84.36 on WMT26 Evaluations
Summary
Cohere and RWS unveil North Small Translate, a 218B-parameter mixture-of-experts system with 25B active parameters whose error-finding Agentic version reaches 84.36 on WMT26 tests and beats DeepL NextGen across every tested non-European region.
Key Points
- North Small Translate uses a 218B-parameter mixture-of-experts architecture with 25B active parameters, supports more than 50 languages, and has 16k-token input and output contexts.
- On WMT26 evaluations, the standard model scores 83.60 across languages and its error-finding Agentic version scores 84.36; both outperform DeepL NextGen in every tested non-European region.
- Cohere develops the model with RWS, and RWS offers it through Language Weaver while Cohere releases weights on Hugging Face for non-commercial research under CC BY-NC 4.0.