الملخص

Recent 4D LiDAR language models aim to reason about objects and their evolving spatial relationships. Yet, in our evaluation, always selecting the same option nearly matches the multiple-choice accuracy of two B4DL-derived configurations. We introduce LiDAR-Hallu, a geometry-referenced benchmark and diagnostic protocol with 10,000 questions across 150 nuScenes scenes. It covers object existence, ego-relative position, distance ordering, relative motion, and temporal localization, with explicit rules for selecting objects, comparing times, and determining reference answers. Our protocol combines fixed-answer and candidate-content controls, cross-scene pairs with identical prompts but opposite reference answers, and relation-specific recall. Analysis of 100,000 recorded responses reveals failures hidden by aggregate accuracy. Candidate duration alone makes temporal answers predictable without observing LiDAR. On paired questions, the models frequently give the same answer to scenes requiring opposite answers. Relation-specific analysis further shows that both configurations miss every positive lateral-motion case across all tested conditions. Temporal-shuffle contrastive decoding provides little net improvement, as repairs are largely offset by new errors and the main failures persist. These results show that evaluating spatio-temporal reasoning requires testing whether models distinguish the queried physical relationships, rather than relying on individual-answer accuracy alone. The source code, checkpoints, and data are released at https://github.com/Awesome4D/4DMLLM_Hallucination_Bench.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Yang, R., Akkoyun, M., Wen, D., Liu, R., Chen, Y., Zheng, J., Wang, X., Yang, K., Paudel, D. P., Van Gool, L., & Peng, K. (2026). Do LiDAR Language Models Really Understand Spatio-temporal Relationships? https://omanscience.com/ar/articles/do-lidar-language-models-really-understand-spatio-temporal-relationships

MLA 9

Yang, Runyi, et al. "Do LiDAR Language Models Really Understand Spatio-temporal Relationships?" https://omanscience.com/ar/articles/do-lidar-language-models-really-understand-spatio-temporal-relationships.

شيكاغو (المؤلف–التاريخ)

Yang, Runyi, Murat Akkoyun, Di Wen, Ruiping Liu, Yufan Chen, Junwei Zheng, Xiaoye Wang, Kailun Yang, Danda Pani Paudel, Luc Van Gool, and Kunyu Peng. 2026. "Do LiDAR Language Models Really Understand Spatio-temporal Relationships?" https://omanscience.com/ar/articles/do-lidar-language-models-really-understand-spatio-temporal-relationships.

هارفارد

Yang, R., Akkoyun, M., Wen, D., Liu, R., Chen, Y., Zheng, J., Wang, X., Yang, K., Paudel, D. P., Van Gool, L. and Peng, K. (2026) 'Do LiDAR Language Models Really Understand Spatio-temporal Relationships?', Available at: https://omanscience.com/ar/articles/do-lidar-language-models-really-understand-spatio-temporal-relationships.

فانكوفر

Yang R, Akkoyun M, Wen D, Liu R, Chen Y, Zheng J, et al. Do LiDAR Language Models Really Understand Spatio-temporal Relationships? https://omanscience.com/ar/articles/do-lidar-language-models-really-understand-spatio-temporal-relationships

IEEE

R. Yang, M. Akkoyun, D. Wen, R. Liu, Y. Chen, J. Zheng, X. Wang, K. Yang, D. P. Paudel, L. Van Gool, and K. Peng, "Do LiDAR Language Models Really Understand Spatio-temporal Relationships?," https://omanscience.com/ar/articles/do-lidar-language-models-really-understand-spatio-temporal-relationships.