Abstract

Autonomous research seeks sustained model improvements through iterative experimentation and feedback. LLM agents show promise in automating machine learning and language-model post-training, but their ability to sustain multimodal improvement remains unclear. We introduce MMPostTrainBench, a benchmark spanning eight tasks in image, audio, video, and joint audio-video understanding and image-grounded software repair. Agents operate from a common base model within fixed budgets, using development feedback before independent evaluation of their submitted models. Evaluation covers target and non-target model outcomes, iterative model improvement and selection, and research integrity. Across all eight tasks, 52.1% of model--task means fall below the base, and evaluated submissions also exhibit non-target regressions. Model performance does not consistently improve across research iterations, and agents do not reliably select the best evaluated candidate for submission; final submissions trail that candidate by up to 5.38 percentage points. Extending autonomous research from text-only to multimodal tasks introduces additional sources of error in perception, cross-modal alignment, and temporal grounding. The observed regressions and selection gaps highlight the need to balance targeted improvements with non-target capability preservation and to retain gains across research iterations. These requirements motivate MMResearch, a multimodal research framework that connects media-grounded evidence to hypotheses and interventions, carries findings across rounds through hierarchical memory, and retains candidates using development evaluation. Added to existing code-agent runtimes, it improves submitted-model accuracy by up to 7.75 percentage points for Claude Opus 4.8 with Claude Code and 2.33 points for GPT-5.6-sol with Codex.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Liu, Y., Wang, Y., Lei, Z., Meng, L., Sun, Y., Lin, J., Liu, H., Chu, Y., Yang, Q., Xu, J., Zhang, L., & Mao, Z. (2026). MMPostTrainBench: Benchmarking Autonomous Research for Multimodal Post-Training. https://omanscience.com/en/articles/mmposttrainbench-benchmarking-autonomous-research-for-multimodal-post-training

MLA 9

Liu, Yuxin, et al. "MMPostTrainBench: Benchmarking Autonomous Research for Multimodal Post-Training." https://omanscience.com/en/articles/mmposttrainbench-benchmarking-autonomous-research-for-multimodal-post-training.

Chicago (author–date)

Liu, Yuxin, Yuxuan Wang, Zhenxin Lei, Lingchen Meng, Yuchong Sun, Junming Lin, Hongcheng Liu, Yunfei Chu, Qize Yang, Jin Xu, Lei Zhang, and Zhendong Mao. 2026. "MMPostTrainBench: Benchmarking Autonomous Research for Multimodal Post-Training." https://omanscience.com/en/articles/mmposttrainbench-benchmarking-autonomous-research-for-multimodal-post-training.

Harvard

Liu, Y., Wang, Y., Lei, Z., Meng, L., Sun, Y., Lin, J., Liu, H., Chu, Y., Yang, Q., Xu, J., Zhang, L. and Mao, Z. (2026) 'MMPostTrainBench: Benchmarking Autonomous Research for Multimodal Post-Training', Available at: https://omanscience.com/en/articles/mmposttrainbench-benchmarking-autonomous-research-for-multimodal-post-training.

Vancouver

Liu Y, Wang Y, Lei Z, Meng L, Sun Y, Lin J, et al. MMPostTrainBench: Benchmarking Autonomous Research for Multimodal Post-Training. https://omanscience.com/en/articles/mmposttrainbench-benchmarking-autonomous-research-for-multimodal-post-training

IEEE

Y. Liu, Y. Wang, Z. Lei, L. Meng, Y. Sun, J. Lin, H. Liu, Y. Chu, Q. Yang, J. Xu, L. Zhang, and Z. Mao, "MMPostTrainBench: Benchmarking Autonomous Research for Multimodal Post-Training," https://omanscience.com/en/articles/mmposttrainbench-benchmarking-autonomous-research-for-multimodal-post-training.