Abstract
Improving joint audio-visual reasoning in Omni Large Language Models typically incurs substantial data construction and training costs. Our diagnostics reveal multi-hop reasoning difficulties despite correct answers to all corresponding single-hop questions and suggest partial decoupling in the local optimization of perception and reasoning objectives. This motivates post-training with different emphases on these capabilities. Text-only reasoning training yields gains across data sources, model scales, and families. With the best-performing text-only configuration, supervised fine-tuning followed by reinforcement learning (RL) raises Qwen2.5-Omni-7B's geometric mean of nine reasoning scores by 25.83% over the base model, outperforming the complete native audio-visual route with 56.6% fewer GPU-hours. Training on data synthesized entirely by a text-only LLM raises this geometric mean by 21.01% without audio-visual data in construction or training. However, text-only training degrades perception. We therefore propose a text-centric post-training paradigm: text-only training provides the main reasoning optimization, and reduced-data native audio-visual RL then refines perception. Refinement uses about 90% fewer input tokens than full-data audio-visual RL, restores perception above the base level, and retains 93.5% of the best-performing text-only pipeline's reasoning gain.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Cheng, Z., Wang, Y., Liu, H., Wu, Q., Fan, J., Qian, C., Wang, Y., & Wang, Y. (2026). Text-Centric Post-Training for Omni-Modal Reasoning. https://omanscience.com/en/articles/text-centric-post-training-for-omni-modal-reasoning
MLA 9
Cheng, Ziyang, et al. "Text-Centric Post-Training for Omni-Modal Reasoning." https://omanscience.com/en/articles/text-centric-post-training-for-omni-modal-reasoning.
Chicago (author–date)
Cheng, Ziyang, Yuhao Wang, Hongcheng Liu, Qimin Wu, Jingru Fan, Chen Qian, Yanfeng Wang, and Yu Wang. 2026. "Text-Centric Post-Training for Omni-Modal Reasoning." https://omanscience.com/en/articles/text-centric-post-training-for-omni-modal-reasoning.
Harvard
Cheng, Z., Wang, Y., Liu, H., Wu, Q., Fan, J., Qian, C., Wang, Y. and Wang, Y. (2026) 'Text-Centric Post-Training for Omni-Modal Reasoning', Available at: https://omanscience.com/en/articles/text-centric-post-training-for-omni-modal-reasoning.
Vancouver
Cheng Z, Wang Y, Liu H, Wu Q, Fan J, Qian C, et al. Text-Centric Post-Training for Omni-Modal Reasoning. https://omanscience.com/en/articles/text-centric-post-training-for-omni-modal-reasoning
IEEE
Z. Cheng, Y. Wang, H. Liu, Q. Wu, J. Fan, C. Qian, Y. Wang, and Y. Wang, "Text-Centric Post-Training for Omni-Modal Reasoning," https://omanscience.com/en/articles/text-centric-post-training-for-omni-modal-reasoning.