Abstract

Robotic operation in previously unseen environments requires both semantic understanding and reliable metric information. While vision--language models (VLMs) provide strong semantic capabilities, their geometric estimates remain less reliable. In this paper, we propose a VLM-driven, modular perception framework for scene understanding using off-the-shelf approaches. Starting from a single RGB-D observation, the scene is segmented into object-level regions, annotated by a VLM, and grounded with depth information to construct a task-independent object-centric representation. Experiments on 151 tabletop scenes show that the proposed decomposition preserves strong semantic performance while substantially improving localization and depth estimation over direct VLM inference. The resulting representation is also integrated with a task-planning framework for robotic execution.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Saccon, E., Faraci, T., Zarzuelo, I. D. L. O., Palopoli, L., Roveri, M., & Saveriano, M. (2026). VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation. https://omanscience.com/en/articles/vlms-can-describe-but-not-measure-object-centric-scene-understanding-for-robotic-manipulation

MLA 9

Saccon, Enrico, et al. "VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation." https://omanscience.com/en/articles/vlms-can-describe-but-not-measure-object-centric-scene-understanding-for-robotic-manipulation.

Chicago (author–date)

Saccon, Enrico, Tommaso Faraci, Iñigo De La Ossa Zarzuelo, Luigi Palopoli, Marco Roveri, and Matteo Saveriano. 2026. "VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation." https://omanscience.com/en/articles/vlms-can-describe-but-not-measure-object-centric-scene-understanding-for-robotic-manipulation.

Harvard

Saccon, E., Faraci, T., Zarzuelo, I. D. L. O., Palopoli, L., Roveri, M. and Saveriano, M. (2026) 'VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation', Available at: https://omanscience.com/en/articles/vlms-can-describe-but-not-measure-object-centric-scene-understanding-for-robotic-manipulation.

Vancouver

Saccon E, Faraci T, Zarzuelo IDLO, Palopoli L, Roveri M, Saveriano M. VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation. https://omanscience.com/en/articles/vlms-can-describe-but-not-measure-object-centric-scene-understanding-for-robotic-manipulation

IEEE

E. Saccon, T. Faraci, I. D. L. O. Zarzuelo, L. Palopoli, M. Roveri, and M. Saveriano, "VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation," https://omanscience.com/en/articles/vlms-can-describe-but-not-measure-object-centric-scene-understanding-for-robotic-manipulation.