Abstract

Large vision-language models (LVLMs) have demonstrated remarkable performance on multimodal reasoning benchmarks, yet their perceptual reliability under physically constrained imaging conditions remains poorly understood. Existing evaluations predominantly assume ideal visual inputs and therefore fail to characterize how camera distance, illumination, viewpoint, and pixel density fundamentally affect semantic recoverability. We introduce SynDORBench, the first physically grounded benchmark for evaluating LVLM perceptual robustness under DORI-calibrated conditions aligned with human visual capability standards. SynDORBench comprises over 54k question--answer pairs generated through a controllable synthetic pipeline that systematically varies viewing distance, lighting, camera geometry, and action pose according to physically interpretable pixel-density regimes. To support scalable low-visibility supervision, we further propose a discernibility annotation framework that propagates human perceptual labels using mask-conditioned statistical features and ensemble learning. We evaluate 16 open-source LVLMs, a commercial LVLM baseline, and YOLO11x across human-presence classification and action recognition tasks under progressively degraded visibility conditions. Our results reveal that perceptual failure in LVLMs is strongly governed by pixel density and physical imaging constraints rather than model scale alone. Surprisingly, several compact open-source LVLMs outperform larger commercial baselines and substantially exceed YOLO11x robustness under long-range and low-light conditions. SynDORBench establishes a new benchmark paradigm for physically grounded multimodal evaluation, enabling systematic analysis of LVLM reliability under real-world perceptual constraints and direct comparison against human visibility thresholds.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Yee, J. S. G., Wang, Z., Zhang, Z., Anand, A., Liu, T., Chan, B., Ng, A. B., & See, S. (2026). SynDORBench: Evaluating LVLM Perceptual Robustness Under Physically Constrained Visibility Conditions. https://omanscience.com/en/articles/syndorbench-evaluating-lvlm-perceptual-robustness-under-physically-constrained-visibility-conditions

MLA 9

Yee, Jeremy Stephen Gabriel, et al. "SynDORBench: Evaluating LVLM Perceptual Robustness Under Physically Constrained Visibility Conditions." https://omanscience.com/en/articles/syndorbench-evaluating-lvlm-perceptual-robustness-under-physically-constrained-visibility-conditions.

Chicago (author–date)

Yee, Jeremy Stephen Gabriel, Zhengkui Wang, Zhiyuan Zhang, Avinash Anand, Timothy Liu, Benedict Chan, Aik Beng Ng, and Simon See. 2026. "SynDORBench: Evaluating LVLM Perceptual Robustness Under Physically Constrained Visibility Conditions." https://omanscience.com/en/articles/syndorbench-evaluating-lvlm-perceptual-robustness-under-physically-constrained-visibility-conditions.

Harvard

Yee, J. S. G., Wang, Z., Zhang, Z., Anand, A., Liu, T., Chan, B., Ng, A. B. and See, S. (2026) 'SynDORBench: Evaluating LVLM Perceptual Robustness Under Physically Constrained Visibility Conditions', Available at: https://omanscience.com/en/articles/syndorbench-evaluating-lvlm-perceptual-robustness-under-physically-constrained-visibility-conditions.

Vancouver

Yee JSG, Wang Z, Zhang Z, Anand A, Liu T, Chan B, et al. SynDORBench: Evaluating LVLM Perceptual Robustness Under Physically Constrained Visibility Conditions. https://omanscience.com/en/articles/syndorbench-evaluating-lvlm-perceptual-robustness-under-physically-constrained-visibility-conditions

IEEE

J. S. G. Yee, Z. Wang, Z. Zhang, A. Anand, T. Liu, B. Chan, A. B. Ng, and S. See, "SynDORBench: Evaluating LVLM Perceptual Robustness Under Physically Constrained Visibility Conditions," https://omanscience.com/en/articles/syndorbench-evaluating-lvlm-perceptual-robustness-under-physically-constrained-visibility-conditions.