Abstract
Detailed image captioning requires accurate and comprehensive descriptions of fine-grained visual content, yet caption quality spans factual accuracy, information coverage, and clarity. Compared with conventional methods that rely mainly on high-quality supervision or holistic rewards, rubric-based reinforcement learning decomposes these requirements into explicit criteria and provides targeted, structured feedback. However, existing methods often use separate models for caption generation, rubric construction, and judging, which may lead to inconsistent interpretations across roles. Some dynamic rubric methods alternate updates between the caption policy and rubric generator while keeping the judge fixed, but staged optimization may still leave rubric construction and judging out of step with policy optimization. We propose MoCo Rubric, a two-stage framework that coordinates these roles. First, role-conditioned, shared-parameter multi-task supervised fine-tuning equips a single vision--language model to serve as the Caption Policy, Rubric Generator, and Rubric Judge. Then, the Generator constructs rubrics online from captions sampled by the current Policy, reference captions, and image evidence. The Judge provides rubric-based rewards, and only the Policy receives GRPO updates. As Policy updates change the candidates being evaluated, we use an exponential moving average of the Policy parameters to update one momentum model shared by the Generator and Judge. This gradual transfer lets both rubric roles track Policy updates without separate RL optimization while smoothing parameter changes that could disrupt their rubric capabilities under direct synchronization. Across five captioning benchmarks, MoCo Rubric achieves an average pairwise win rate of 72.83\%, the best mean rank in blind ranking, and the highest average score in caption-based question answering.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Ji, Z., Jin, L., Wang, S., Lu, J., Lu, C., Wu, Y., Hu, Y., Cui, L., & Xu, Y. (2026). Momentum-Coupled Rubric Adaptation for Detailed Image Captioning. https://omanscience.com/en/articles/momentum-coupled-rubric-adaptation-for-detailed-image-captioning
MLA 9
Ji, Zhenwen, et al. "Momentum-Coupled Rubric Adaptation for Detailed Image Captioning." https://omanscience.com/en/articles/momentum-coupled-rubric-adaptation-for-detailed-image-captioning.
Chicago (author–date)
Ji, Zhenwen, Lei Jin, Shanyong Wang, Jiaming Lu, Chengqiang Lu, Yi Wu, Yao Hu, Lizhen Cui, and Yanyu Xu. 2026. "Momentum-Coupled Rubric Adaptation for Detailed Image Captioning." https://omanscience.com/en/articles/momentum-coupled-rubric-adaptation-for-detailed-image-captioning.
Harvard
Ji, Z., Jin, L., Wang, S., Lu, J., Lu, C., Wu, Y., Hu, Y., Cui, L. and Xu, Y. (2026) 'Momentum-Coupled Rubric Adaptation for Detailed Image Captioning', Available at: https://omanscience.com/en/articles/momentum-coupled-rubric-adaptation-for-detailed-image-captioning.
Vancouver
Ji Z, Jin L, Wang S, Lu J, Lu C, Wu Y, et al. Momentum-Coupled Rubric Adaptation for Detailed Image Captioning. https://omanscience.com/en/articles/momentum-coupled-rubric-adaptation-for-detailed-image-captioning
IEEE
Z. Ji, L. Jin, S. Wang, J. Lu, C. Lu, Y. Wu, Y. Hu, L. Cui, and Y. Xu, "Momentum-Coupled Rubric Adaptation for Detailed Image Captioning," https://omanscience.com/en/articles/momentum-coupled-rubric-adaptation-for-detailed-image-captioning.