الملخص

Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to act toward a goal and anticipates the resulting scene changes. We introduce ChronoGraph, a functional 4D scene graph that links actions on affordance parts to semantic and geometric state changes. By representing observed and anticipated transitions in the same form, it provides a shared basis for understanding and planning. We construct ChronoGraphBench through an automatic data engine that converts human-interaction videos and simulated robot trajectories into graph-annotated questions for training and evaluating Vision-Language Models (VLMs) on both tasks. Using these annotations, we train ChronoGraphVLM by adapting pretrained VLMs in two stages. Graph-as-Chain-of-Thought supervised fine-tuning teaches the models to reconstruct observed transitions and predict future ones as graph traces before answering. Subsequent joint 4D graph reinforcement learning directly rewards graph properties and answer correctness. Experiments across model scales show improvements over the corresponding pretrained baselines and zero-shot transfer to VLM4D. Real-world demonstrations further show that graph-based planning and affordance grounding support mobile manipulation through existing robot skills without additional fine-tuning.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Zhang, C., Gwiazda, M., Jiao, G., Ju, Y., Tombari, F., Sreenath, K., Pollefeys, M., & Hong, S. (2026). ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning. https://omanscience.com/ar/articles/chronograph-functional-4d-scene-graphs-with-vision-language-models-for-interaction-understanding-and-grounded-planning

MLA 9

Zhang, Chenyangguang, et al. "ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning." https://omanscience.com/ar/articles/chronograph-functional-4d-scene-graphs-with-vision-language-models-for-interaction-understanding-and-grounded-planning.

شيكاغو (المؤلف–التاريخ)

Zhang, Chenyangguang, Malgorzata Gwiazda, Guanlong Jiao, Yuanchen Ju, Federico Tombari, Koushil Sreenath, Marc Pollefeys, and Sunghwan Hong. 2026. "ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning." https://omanscience.com/ar/articles/chronograph-functional-4d-scene-graphs-with-vision-language-models-for-interaction-understanding-and-grounded-planning.

هارفارد

Zhang, C., Gwiazda, M., Jiao, G., Ju, Y., Tombari, F., Sreenath, K., Pollefeys, M. and Hong, S. (2026) 'ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning', Available at: https://omanscience.com/ar/articles/chronograph-functional-4d-scene-graphs-with-vision-language-models-for-interaction-understanding-and-grounded-planning.

فانكوفر

Zhang C, Gwiazda M, Jiao G, Ju Y, Tombari F, Sreenath K, et al. ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning. https://omanscience.com/ar/articles/chronograph-functional-4d-scene-graphs-with-vision-language-models-for-interaction-understanding-and-grounded-planning

IEEE

C. Zhang, M. Gwiazda, G. Jiao, Y. Ju, F. Tombari, K. Sreenath, M. Pollefeys, and S. Hong, "ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning," https://omanscience.com/ar/articles/chronograph-functional-4d-scene-graphs-with-vision-language-models-for-interaction-understanding-and-grounded-planning.