الملخص
Existing geospatial vision-language models (Geo-VLMs) typically optimize diverse geospatial tasks through a unified multi-task adaptation paradigm without explicitly accounting for the heterogeneous optimization characteristics. Our empirical observations reveal heterogeneous gradient characteristics across tasks, including vision-language differences, intra-branch gradient relationships, and task interference, which hinder effective multi-task optimization. Motivated by these observations, we propose Gradient-Guided Decoupled Adaptation (G2DA), a gradient-aware optimization framework for multi-task Geo-VLM learning. G2DA first partitions tasks into vision- and language-centric groups through gradient-guided cross-modal decoupling. It then constructs modality-specific curricula based on task gradient similarity and employs bidirectional rehearsal to mitigate the recency effects introduced by sequential optimization. We evaluate G2DA on three Geo-VLM benchmarks using six InternVL3 and Qwen3.5-VL variants, along with GeoChat and GeoLLaVA. Across all 24 benchmark-model combinations, G2DA consistently outperforms representative baselines, improving over the strongest competitor by 3.08, 4.30, and 2.81 percentage points on UrBench-MCQ, XLRS-Bench-Lite, and VRS-Bench-VQA, respectively. These results demonstrate the effectiveness of gradient-guided task organization for Geo-VLM adaptation.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Wang, D., Balakrishnan, D., Srinivasan, R., & Wang, S. (2026). Gradient-Guided Decoupled Adaptation for Geospatial Vision-Language Models. https://omanscience.com/ar/articles/gradient-guided-decoupled-adaptation-for-geospatial-vision-language-models
MLA 9
Wang, Dongdong, et al. "Gradient-Guided Decoupled Adaptation for Geospatial Vision-Language Models." https://omanscience.com/ar/articles/gradient-guided-decoupled-adaptation-for-geospatial-vision-language-models.
شيكاغو (المؤلف–التاريخ)
Wang, Dongdong, Deepak Balakrishnan, Ravi Srinivasan, and Shenhao Wang. 2026. "Gradient-Guided Decoupled Adaptation for Geospatial Vision-Language Models." https://omanscience.com/ar/articles/gradient-guided-decoupled-adaptation-for-geospatial-vision-language-models.
هارفارد
Wang, D., Balakrishnan, D., Srinivasan, R. and Wang, S. (2026) 'Gradient-Guided Decoupled Adaptation for Geospatial Vision-Language Models', Available at: https://omanscience.com/ar/articles/gradient-guided-decoupled-adaptation-for-geospatial-vision-language-models.
فانكوفر
Wang D, Balakrishnan D, Srinivasan R, Wang S. Gradient-Guided Decoupled Adaptation for Geospatial Vision-Language Models. https://omanscience.com/ar/articles/gradient-guided-decoupled-adaptation-for-geospatial-vision-language-models
IEEE
D. Wang, D. Balakrishnan, R. Srinivasan, and S. Wang, "Gradient-Guided Decoupled Adaptation for Geospatial Vision-Language Models," https://omanscience.com/ar/articles/gradient-guided-decoupled-adaptation-for-geospatial-vision-language-models.