الملخص
Unified multimodal models (UMMs) understand and generate both text and images, which lets a model produce its own training data. Existing self-improvement in UMMs keeps supervision on the visual side, where image understanding judges image generation. We propose recursive cross-capability self-improvement (RSI), a training loop in which the text and visual abilities of a UMM supply training data for one another. In each round, the model generates images and reads them to find where it falls short. It then writes programs aimed at these shortcomings, and execution verifies every result against its specification. Verified renders train image generation, while labeled renders and the model's own correct programs train visual understanding and program writing. Program execution thus acts as a source of truth outside the model, so errors do not accumulate across rounds. We study RSI on charts and build BasicChartBench to evaluate open models early in training. On requests worded differently from training, four rounds of RSI raise the score from 45.7% to 60.2%, while continued training stays at 46.3%. Verified construction carries most of the gain, and targeting the model's failures adds 3.5%. Along the way, the share of verified programs rises from 48.9% to 95.2%, and the reader's accuracy on edited renders rises from 55.6% to 87.4%.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Wang, H., Shi, C., Yang, C., Wu, Y., Berg-Kirkpatrick, T., & Ma, X. (2026). Recursive Self-Improvement in Unified Multimodal Models. https://omanscience.com/ar/articles/recursive-self-improvement-in-unified-multimodal-models
MLA 9
Wang, Huijuan, et al. "Recursive Self-Improvement in Unified Multimodal Models." https://omanscience.com/ar/articles/recursive-self-improvement-in-unified-multimodal-models.
شيكاغو (المؤلف–التاريخ)
Wang, Huijuan, Chufan Shi, Cheng Yang, Yaokang Wu, Taylor Berg-Kirkpatrick, and Xuezhe Ma. 2026. "Recursive Self-Improvement in Unified Multimodal Models." https://omanscience.com/ar/articles/recursive-self-improvement-in-unified-multimodal-models.
هارفارد
Wang, H., Shi, C., Yang, C., Wu, Y., Berg-Kirkpatrick, T. and Ma, X. (2026) 'Recursive Self-Improvement in Unified Multimodal Models', Available at: https://omanscience.com/ar/articles/recursive-self-improvement-in-unified-multimodal-models.
فانكوفر
Wang H, Shi C, Yang C, Wu Y, Berg-Kirkpatrick T, Ma X. Recursive Self-Improvement in Unified Multimodal Models. https://omanscience.com/ar/articles/recursive-self-improvement-in-unified-multimodal-models
IEEE
H. Wang, C. Shi, C. Yang, Y. Wu, T. Berg-Kirkpatrick, and X. Ma, "Recursive Self-Improvement in Unified Multimodal Models," https://omanscience.com/ar/articles/recursive-self-improvement-in-unified-multimodal-models.