الملخص

Pathological assessment relies on recognizing fine-grained visual details in histological images. Vision-language models (VLMs) increasingly support pathology interpretation, yet their ability to perceive these details remains inadequate. This weakness leads to inaccurate cellular observations that can persist even when final answers are correct. In this paper, we propose ASPECT to improve visually grounded reasoning through explicit supervision of cellular appearance and abundance. ASPECT trains intermediate visual tokens through pathology feature reconstruction, cell feature alignment, and count supervision. Three-stage supervised fine-tuning teaches the model to perceive, generate visual tokens, and reason, followed by reinforcement learning that rewards answer correctness and consistency with reported measurements. We also introduce PathoVernier, a benchmark of 759 expert-reviewed questions from five pathology datasets covering four cellular composition tasks. It evaluates both final answers and intermediate measurements to expose errors hidden by answer accuracy. On PathoVernier, ASPECT achieves relative accuracy gains of approximately 19.2% over the strongest baseline, Gemini-3.1-Pro, and 99.3% over its Qwen3-VL-8B backbone, while reducing RAWR, which measures counting errors within correct responses, by 28.1% and 42.7%, respectively. ASPECT also improves over its backbone on three external pathology benchmarks covering classification and question answering beyond cellular composition tasks.

الكلمات المفتاحية

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Zhang, C., Zhang, W., Li, B., Li, M., Liu, X., Yang, J., Chen, J., Zhang, Z., Yi, Y., Bu, H., & Lv, J. (2026). See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology. https://omanscience.com/ar/articles/see-measure-and-reason-learning-visually-grounded-reasoning-in-pathology

MLA 9

Zhang, Chengyang, et al. "See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology." https://omanscience.com/ar/articles/see-measure-and-reason-learning-visually-grounded-reasoning-in-pathology.

شيكاغو (المؤلف–التاريخ)

Zhang, Chengyang, Wenchuan Zhang, Bo Li, Mengran Li, Xinyu Liu, Jiaming Yang, Jie Chen, Zhang Zhang, Yuhao Yi, Hong Bu, and Jiancheng Lv. 2026. "See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology." https://omanscience.com/ar/articles/see-measure-and-reason-learning-visually-grounded-reasoning-in-pathology.

هارفارد

Zhang, C., Zhang, W., Li, B., Li, M., Liu, X., Yang, J., Chen, J., Zhang, Z., Yi, Y., Bu, H. and Lv, J. (2026) 'See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology', Available at: https://omanscience.com/ar/articles/see-measure-and-reason-learning-visually-grounded-reasoning-in-pathology.

فانكوفر

Zhang C, Zhang W, Li B, Li M, Liu X, Yang J, et al. See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology. https://omanscience.com/ar/articles/see-measure-and-reason-learning-visually-grounded-reasoning-in-pathology

IEEE

C. Zhang, W. Zhang, B. Li, M. Li, X. Liu, J. Yang, J. Chen, Z. Zhang, Y. Yi, H. Bu, and J. Lv, "See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology," https://omanscience.com/ar/articles/see-measure-and-reason-learning-visually-grounded-reasoning-in-pathology.