Cohere Labs has published research on how to make reasoning models perform their intermediate reasoning in the same language as a user's prompt rather than defaulting to English. The work introduces Tiny Aya L2-Thinker, a 3.35-billion-parameter research model, and argues that careful composition of supervised fine-tuning data can transfer in-language reasoning to many languages without requiring large reasoning datasets for every language.
The training strategy combines three kinds of data. English reasoning examples provide a general problem-solving backbone. A much smaller amount of multilingual reasoning data teaches the model to express reasoning in additional languages. Multilingual non-reasoning question-and-answer data then helps that behavior generalize, including to languages for which the model did not receive dedicated reasoning examples. Cohere reports that adding multilingual reasoning data increased the in-language reasoning rate sharply, while the final mixture further improved both generalization and task accuracy.
The researchers evaluated the model across six benchmarks and 60 languages, covering mathematics, commonsense and cultural reasoning, instruction following and open-ended generation. Cohere reports that Tiny Aya L2-Thinker produced reasoning in the prompt language more than 93% of the time while maintaining similar accuracy to a matched English-reasoning model on most benchmarks. The paper also identifies an important exception: performance falls more noticeably on PolyMath, a harder competition-level mathematics benchmark, showing that language alignment does not remove all quality trade-offs.
The study also examines inference-time attempts to force another model to reason in the target language. According to Cohere, prompting Qwen3.5-4B to do so had limited effect, while prefilling the beginning of its reasoning in the target language improved language compliance at an accuracy cost. Tiny Aya L2-Thinker was also reported to use fewer reasoning tokens and exhibit less repetitive 'doomlooping' than the compared configuration. These results are research claims from Cohere and the associated paper rather than independent replication.
The work is relevant because multilingual AI quality is often measured only by the language of the final answer. If a model internally reasons in English and then returns a translated response, users may be unable to inspect the reasoning in their own language, and language-specific or cultural information may be handled less naturally. Cohere has released the Tiny Aya L2-Thinker weights and multilingual reasoning data, making the work more inspectable and reproducible. However, this run found no credible independent replication of the reported 93% result. The strongest conclusion is therefore that Cohere has presented and openly released a concrete training recipe and research model for in-language reasoning, while the breadth of its claimed generalization still needs external testing.