[
    {
        "id": "osp-20623",
        "type": "article-journal",
        "title": "What Does Post-Training Change in Multilingual Reasoning?",
        "author": [
            {
                "family": "Li",
                "given": "Hongyang"
            },
            {
                "family": "Li",
                "given": "Xiao"
            },
            {
                "family": "Wu",
                "given": "Caesar"
            },
            {
                "family": "Danoy",
                "given": "Grégoire"
            },
            {
                "family": "Bouvry",
                "given": "Pascal"
            }
        ],
        "URL": "https://omanscience.com/ar/articles/what-does-post-training-change-in-multilingual-reasoning",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Open-source reasoning models provide unequal access to reasoning capability across languages. When a model can solve a problem but cannot deliver a complete solution in the user's language, language becomes an access barrier rather than merely a source of performance variation. We audit Qwen3 checkpoints on competition-mathematics tasks in eleven languages. Across the ten non-English languages, only 15.4-17.9% of problems receive a correct, terminating solution with visible reasoning in the requested language in any of 16 samples, compared with 92.9% in English. To identify the source of this disparity, we evaluate thirteen endpoints from one model family, spanning released checkpoints, multilingual supervised fine-tuning (SFT) at two scales, controlled SFT ablations, and three reinforcement-learning (RL) reward formulations. We jointly track correctness, language adherence, termination, and delivery efficiency. The dominant bottleneck shifts across post-training stages. Released models often reason in English. Multilingual SFT restores target-language reasoning, but accuracy declines across multilingual, English-only, and single-language SFT runs, showing that this cost is not specific to multilingual mixing; non-English reasoning traces additionally become prone to non-terminating loops. RL restores termination in both arms at no cost in accuracy, but only the arm whose reward includes a language term delivers: rewarding correctness alone returns the model to English. Together, these stages establish a constructive post-training path from English-pivoted capability to multilingual reasoning that is reliably delivered."
    }
]