الملخص

Urban diagnosis integrates heterogeneous observations to identify urban problems, localize affected areas, and investigate contributing factors, informing evidence-based urban planning and management. However, its reliance on labor-intensive, case-specific expert workflows limits scalability and reuse, motivating the exploration of agent-based execution. To evaluate this capability, we introduce DUDA-Bench, a hierarchical and interactive benchmark that formalizes data-driven urban diagnosis as a multi-stage agent workflow. It comprises 86 atomic and 22 workflow tasks spanning four analytical stages, grounded in multimodal data from 12 cities covering five urban problem types. Evaluations of seven backbone models and five agent systems reveal a substantial gap between isolated analytical competence and end-to-end diagnosis, with system benefits varying across backbones. Trajectory analysis shows that unresolved evidence gaps propagate across stages, while successful recovery involves revising assumptions and actions using feedback. These findings highlight limitations in coordinating analytical capabilities across stages, particularly adaptive planning, evidence integration, and verification. More broadly, DUDA-Bench provides a framework for translating expert analytical workflows into hierarchical agent tasks and process-aware evaluation, supporting systematic assessment of end-to-end analytical capabilities.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Song, Y., Ni, H., Zhang, W., & Liu, H. (2026). DUDA-Bench: Benchmarking LLM Agents on Multimodal Data-Driven Urban Diagnosis. https://omanscience.com/ar/articles/duda-bench-benchmarking-llm-agents-on-multimodal-data-driven-urban-diagnosis

MLA 9

Song, Yizhi, et al. "DUDA-Bench: Benchmarking LLM Agents on Multimodal Data-Driven Urban Diagnosis." https://omanscience.com/ar/articles/duda-bench-benchmarking-llm-agents-on-multimodal-data-driven-urban-diagnosis.

شيكاغو (المؤلف–التاريخ)

Song, Yizhi, Hang Ni, Weijia Zhang, and Hao Liu. 2026. "DUDA-Bench: Benchmarking LLM Agents on Multimodal Data-Driven Urban Diagnosis." https://omanscience.com/ar/articles/duda-bench-benchmarking-llm-agents-on-multimodal-data-driven-urban-diagnosis.

هارفارد

Song, Y., Ni, H., Zhang, W. and Liu, H. (2026) 'DUDA-Bench: Benchmarking LLM Agents on Multimodal Data-Driven Urban Diagnosis', Available at: https://omanscience.com/ar/articles/duda-bench-benchmarking-llm-agents-on-multimodal-data-driven-urban-diagnosis.

فانكوفر

Song Y, Ni H, Zhang W, Liu H. DUDA-Bench: Benchmarking LLM Agents on Multimodal Data-Driven Urban Diagnosis. https://omanscience.com/ar/articles/duda-bench-benchmarking-llm-agents-on-multimodal-data-driven-urban-diagnosis

IEEE

Y. Song, H. Ni, W. Zhang, and H. Liu, "DUDA-Bench: Benchmarking LLM Agents on Multimodal Data-Driven Urban Diagnosis," https://omanscience.com/ar/articles/duda-bench-benchmarking-llm-agents-on-multimodal-data-driven-urban-diagnosis.