الملخص

High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover detail, but it asks the main model to choose a region before obtaining a reliable overview. We introduce VPS, a visual parallel-search framework in which a main agent first invokes grid_search to inspect image tiles in parallel with question-conditioned sub-agents, and then adaptively invokes zoom_in on a precise or merged region. The same interface supports both training-free inference and post-training of the main and sub-agents. Across five benchmark splits and three model sizes, VPS improves mean accuracy over dedicated zoom-only search in 14 of 15 same-model comparisons, with gains up to 8.0 points and especially strong improvements for smaller main models. ZoomBench retains an approximately 3.2-point gain at every tested size. We further develop a supervision pipeline with hint-free verification and a paired role-specific GRPO surrogate for learning the controller and tile-reader roles. SFT improves observed accuracy on all five benchmark splits, including a 4.17-point gain on HR-Bench 4K. Role-specific RL further reshapes search behavior: main-only RL reduces mean tool use from 2.65 to 2.11 with similar pass@1 in an internal four-response evaluation, while external accuracy changes are mixed. Joint training reveals an asymmetry between local evidence reading and global search control. Together, these results support VPS as an effective inference-time scaffold and a trainable decomposition for visual evidence acquisition.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Tao, X., Teng, Y., Fu, X., Gong, C., Liu, Z., Xie, X., Liu, R., & Kong, L. (2026). Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom. https://omanscience.com/ar/articles/visual-parallel-search-learning-to-search-high-resolution-images-with-parallel-tile-inspection-and-adaptive-zoom

MLA 9

Tao, Xijia, et al. "Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom." https://omanscience.com/ar/articles/visual-parallel-search-learning-to-search-high-resolution-images-with-parallel-tile-inspection-and-adaptive-zoom.

شيكاغو (المؤلف–التاريخ)

Tao, Xijia, Yihua Teng, Xinyu Fu, Cheng Gong, Ziru Liu, Xudong Xie, Rui Liu, and Lingpeng Kong. 2026. "Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom." https://omanscience.com/ar/articles/visual-parallel-search-learning-to-search-high-resolution-images-with-parallel-tile-inspection-and-adaptive-zoom.

هارفارد

Tao, X., Teng, Y., Fu, X., Gong, C., Liu, Z., Xie, X., Liu, R. and Kong, L. (2026) 'Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom', Available at: https://omanscience.com/ar/articles/visual-parallel-search-learning-to-search-high-resolution-images-with-parallel-tile-inspection-and-adaptive-zoom.

فانكوفر

Tao X, Teng Y, Fu X, Gong C, Liu Z, Xie X, et al. Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom. https://omanscience.com/ar/articles/visual-parallel-search-learning-to-search-high-resolution-images-with-parallel-tile-inspection-and-adaptive-zoom

IEEE

X. Tao, Y. Teng, X. Fu, C. Gong, Z. Liu, X. Xie, R. Liu, and L. Kong, "Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom," https://omanscience.com/ar/articles/visual-parallel-search-learning-to-search-high-resolution-images-with-parallel-tile-inspection-and-adaptive-zoom.