[
    {
        "id": "osp-20650",
        "type": "article-journal",
        "title": "Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom",
        "author": [
            {
                "family": "Tao",
                "given": "Xijia"
            },
            {
                "family": "Teng",
                "given": "Yihua"
            },
            {
                "family": "Fu",
                "given": "Xinyu"
            },
            {
                "family": "Gong",
                "given": "Cheng"
            },
            {
                "family": "Liu",
                "given": "Ziru"
            },
            {
                "family": "Xie",
                "given": "Xudong"
            },
            {
                "family": "Liu",
                "given": "Rui"
            },
            {
                "family": "Kong",
                "given": "Lingpeng"
            }
        ],
        "URL": "https://omanscience.com/en/articles/visual-parallel-search-learning-to-search-high-resolution-images-with-parallel-tile-inspection-and-adaptive-zoom",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover detail, but it asks the main model to choose a region before obtaining a reliable overview. We introduce VPS, a visual parallel-search framework in which a main agent first invokes grid_search to inspect image tiles in parallel with question-conditioned sub-agents, and then adaptively invokes zoom_in on a precise or merged region. The same interface supports both training-free inference and post-training of the main and sub-agents. Across five benchmark splits and three model sizes, VPS improves mean accuracy over dedicated zoom-only search in 14 of 15 same-model comparisons, with gains up to 8.0 points and especially strong improvements for smaller main models. ZoomBench retains an approximately 3.2-point gain at every tested size. We further develop a supervision pipeline with hint-free verification and a paired role-specific GRPO surrogate for learning the controller and tile-reader roles. SFT improves observed accuracy on all five benchmark splits, including a 4.17-point gain on HR-Bench 4K. Role-specific RL further reshapes search behavior: main-only RL reduces mean tool use from 2.65 to 2.11 with similar pass@1 in an internal four-response evaluation, while external accuracy changes are mixed. Joint training reveals an asymmetry between local evidence reading and global search control. Together, these results support VPS as an effective inference-time scaffold and a trainable decomposition for visual evidence acquisition."
    }
]