الباحثون

Shenhao Wang

المنشورات 2

نسخة أولية وصول مفتوح

Gradient-Guided Decoupled Adaptation for Geospatial Vision-Language Models

Existing geospatial vision-language models (Geo-VLMs) typically optimize diverse geospatial tasks through a unified multi-task adaptation paradigm without explicitly accounting for the heterogeneous optimization characteristics. Our empirical observations reveal heterogeneous gradient characteristics across tasks, incl …

نسخة أولية وصول مفتوح

USAI-Quant: A Quantitative Reasoning Benchmark for Vision-Language Models in Built Environments

Dongdong Wang, Qingqi Song, Yuzhou Chen وآخرون · 2026

Large vision-language models (VLMs) have emerged as a powerful paradigm for urban and spatial AI. However, current state-of-the-art large VLMs still struggle with quantitative reasoning on remote sensing imagery. Existing benchmarks and algorithms are predominantly based on qualitative Visual Question Answering (VQA), …

المؤلفون المشاركون