Abstract
Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preservation across languages. However, existing benchmarks have limited language coverage and often lack explicit image-specific evaluation criteria, making it difficult to comprehensively assess this capability. To systematically evaluate this capability, we introduce VISTA-Bench, covering 22 languages and 10 domains, and develop an image-specific rubric evaluation protocol. The benchmark combines sampling for language and scenario coverage with model-assisted, human-verified annotations that group related text into coherent semantic units and provide multilingual reference translations. The rubrics specify essential content, semantic relations, and acceptable translation variants, yielding separate output-based scores for translation quality and the preservation of visual and knowledge-dependent information. We conduct extensive evaluations of 16 mainstream models, including 12 multimodal models and four text-input models, and provide systematic analyses across languages, domains, and evaluation dimensions.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Lv, B., Zheng, M., Li, Z., Liu, F., Sun, M., & Chen, T. (2026). VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics. https://omanscience.com/en/articles/vista-bench-benchmarking-multilingual-image-translation-with-image-specific-rubrics
MLA 9
Lv, Bo, et al. "VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics." https://omanscience.com/en/articles/vista-bench-benchmarking-multilingual-image-translation-with-image-specific-rubrics.
Chicago (author–date)
Lv, Bo, Mao Zheng, Zheng Li, Fangxu Liu, Mingrui Sun, and Tao Chen. 2026. "VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics." https://omanscience.com/en/articles/vista-bench-benchmarking-multilingual-image-translation-with-image-specific-rubrics.
Harvard
Lv, B., Zheng, M., Li, Z., Liu, F., Sun, M. and Chen, T. (2026) 'VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics', Available at: https://omanscience.com/en/articles/vista-bench-benchmarking-multilingual-image-translation-with-image-specific-rubrics.
Vancouver
Lv B, Zheng M, Li Z, Liu F, Sun M, Chen T. VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics. https://omanscience.com/en/articles/vista-bench-benchmarking-multilingual-image-translation-with-image-specific-rubrics
IEEE
B. Lv, M. Zheng, Z. Li, F. Liu, M. Sun, and T. Chen, "VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics," https://omanscience.com/en/articles/vista-bench-benchmarking-multilingual-image-translation-with-image-specific-rubrics.