Abstract

Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preservation across languages. However, existing benchmarks have limited language coverage and often lack explicit image-specific evaluation criteria, making it difficult to comprehensively assess this capability. To systematically evaluate this capability, we introduce VISTA-Bench, covering 22 languages and 10 domains, and develop an image-specific rubric evaluation protocol. The benchmark combines sampling for language and scenario coverage with model-assisted, human-verified annotations that group related text into coherent semantic units and provide multilingual reference translations. The rubrics specify essential content, semantic relations, and acceptable translation variants, yielding separate output-based scores for translation quality and the preservation of visual and knowledge-dependent information. We conduct extensive evaluations of 16 mainstream models, including 12 multimodal models and four text-input models, and provide systematic analyses across languages, domains, and evaluation dimensions.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Lv, B., Zheng, M., Li, Z., Liu, F., Sun, M., & Chen, T. (2026). VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics. https://omanscience.com/en/articles/vista-bench-benchmarking-multilingual-image-translation-with-image-specific-rubrics

MLA 9

Lv, Bo, et al. "VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics." https://omanscience.com/en/articles/vista-bench-benchmarking-multilingual-image-translation-with-image-specific-rubrics.

Chicago (author–date)

Lv, Bo, Mao Zheng, Zheng Li, Fangxu Liu, Mingrui Sun, and Tao Chen. 2026. "VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics." https://omanscience.com/en/articles/vista-bench-benchmarking-multilingual-image-translation-with-image-specific-rubrics.

Harvard

Lv, B., Zheng, M., Li, Z., Liu, F., Sun, M. and Chen, T. (2026) 'VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics', Available at: https://omanscience.com/en/articles/vista-bench-benchmarking-multilingual-image-translation-with-image-specific-rubrics.

Vancouver

Lv B, Zheng M, Li Z, Liu F, Sun M, Chen T. VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics. https://omanscience.com/en/articles/vista-bench-benchmarking-multilingual-image-translation-with-image-specific-rubrics

IEEE

B. Lv, M. Zheng, Z. Li, F. Liu, M. Sun, and T. Chen, "VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics," https://omanscience.com/en/articles/vista-bench-benchmarking-multilingual-image-translation-with-image-specific-rubrics.