Abstract

Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in ultrasound image understanding. However, these models fail to provide pixel-level visual evidence aligned with their semantic predictions, and their fine-grained grounding capability in ultrasound remains largely unclear. We introduce UltraG-Bench, a large-scale multi-task benchmark for evaluating pixel-level evidence grounding in ultrasound. UltraG-Bench is built by annotating 40 public ultrasound segmentation datasets spanning 13 anatomical categories, and comprises three progressive tasks: instruction-guided segmentation, evidence-grounded VQA, and evidence-grounded report generation, with 331125, 666779, and 138832 annotations, respectively. Comprehensive evaluation of 14 state-of-the-art models reveals a substantial gap between semantic understanding and fine-grained pixel-level localization. We further propose UltraG-Agent, which combines the semantic reasoning capabilities of a VLM with the ultrasound-specific segmentation capability of UltraSAM3. Experiments show that UltraG-Agent substantially improves both semantic prediction and pixel-level visual grounding. Our dataset and code are available at https://github.com/zhuqh19/UltraG-Bench.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhu, Q., Xu, B., Lin, R., Wang, C., Shao, Y., Zhu, B., Sun, J., Zhao, L., Lin, H., & Xia, F. (2026). UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound. https://omanscience.com/en/articles/ultrag-bench-a-multi-task-benchmark-for-assessing-large-vision-language-models-on-pixel-level-evidence-grounding-in-ultrasound

MLA 9

Zhu, Quanhao, et al. "UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound." https://omanscience.com/en/articles/ultrag-bench-a-multi-task-benchmark-for-assessing-large-vision-language-models-on-pixel-level-evidence-grounding-in-ultrasound.

Chicago (author–date)

Zhu, Quanhao, Bo Xu, Rui Lin, Chenyuan Wang, Yu Shao, Boling Zhu, Jiuyan Sun, Liang Zhao, Hongfei Lin, and Feng Xia. 2026. "UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound." https://omanscience.com/en/articles/ultrag-bench-a-multi-task-benchmark-for-assessing-large-vision-language-models-on-pixel-level-evidence-grounding-in-ultrasound.

Harvard

Zhu, Q., Xu, B., Lin, R., Wang, C., Shao, Y., Zhu, B., Sun, J., Zhao, L., Lin, H. and Xia, F. (2026) 'UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound', Available at: https://omanscience.com/en/articles/ultrag-bench-a-multi-task-benchmark-for-assessing-large-vision-language-models-on-pixel-level-evidence-grounding-in-ultrasound.

Vancouver

Zhu Q, Xu B, Lin R, Wang C, Shao Y, Zhu B, et al. UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound. https://omanscience.com/en/articles/ultrag-bench-a-multi-task-benchmark-for-assessing-large-vision-language-models-on-pixel-level-evidence-grounding-in-ultrasound

IEEE

Q. Zhu, B. Xu, R. Lin, C. Wang, Y. Shao, B. Zhu, J. Sun, L. Zhao, H. Lin, and F. Xia, "UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound," https://omanscience.com/en/articles/ultrag-bench-a-multi-task-benchmark-for-assessing-large-vision-language-models-on-pixel-level-evidence-grounding-in-ultrasound.