Abstract

Monocular metric depth estimation and 3D visual grounding represent the two complementary cornerstones of monocular 3D spatial understanding (M3Sun), from which the fundamental 3D spatial information required by M3Sun can be acquired. However, these complementary tasks are generally conducted by separate frameworks, which pose challenges of inflexible and unaligned spatial information access for embodied intelligence systems. In this paper, we propose a unified agent for monocular 3D spatial understanding (M3SunAgent) that leverages a large language model (LLM) as a task planner for spatial visual programming, which flexibly generate structured programs and coordinate tools. For instance-level metric depth estimation task, M3SunAgent invokes an object detector tool to locate the target, estimates depth at selected points with a depth estimation tool, and aggregates these predictions into an instance-level depth estimate. We also construct the M3Sun Instance (M3SI) dataset, a benchmark with 2,910 samples for evaluation. For monocular 3D visual grounding task, M3SunAgent uses a vision-language model (VLM) tool to locate the target and output basic spatial attributes, then combines back-projection tool with a dimension-lifting tool to predict its 3D bounding box. Experimental results demonstrate the superior performance of M3SunAgent. Specifically, in evaluations of instance-level monocular metric depth estimation, M3SunAgent achieves the best performance among all compared models, 52.61% of predicted instances are distributed below depth error 0.25 ($δ< 0.25$). In evaluations of monocular 3D visual grounding, M3SunAgent demonstrates overall competitive performance than vision and VLM models, reaching a 3D mean intersection over union (mIoU) of 41.73% and exceeding the state-of-the-art MonoVLM model by 3.62%.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhang, J., Wu, K., Zhu, M., Qiao, R., Cai, C., & Li, Z. (2026). M3SunAgent: Monocular 3D Spatial Understanding Agent for Metric Depth Estimation and 3D Visual Grounding. https://omanscience.com/en/articles/m3sunagent-monocular-3d-spatial-understanding-agent-for-metric-depth-estimation-and-3d-visual-grounding

MLA 9

Zhang, Jinsong, et al. "M3SunAgent: Monocular 3D Spatial Understanding Agent for Metric Depth Estimation and 3D Visual Grounding." https://omanscience.com/en/articles/m3sunagent-monocular-3d-spatial-understanding-agent-for-metric-depth-estimation-and-3d-visual-grounding.

Chicago (author–date)

Zhang, Jinsong, Kejun Wu, Ming Zhu, Renjie Qiao, Chengtao Cai, and Zhengguo Li. 2026. "M3SunAgent: Monocular 3D Spatial Understanding Agent for Metric Depth Estimation and 3D Visual Grounding." https://omanscience.com/en/articles/m3sunagent-monocular-3d-spatial-understanding-agent-for-metric-depth-estimation-and-3d-visual-grounding.

Harvard

Zhang, J., Wu, K., Zhu, M., Qiao, R., Cai, C. and Li, Z. (2026) 'M3SunAgent: Monocular 3D Spatial Understanding Agent for Metric Depth Estimation and 3D Visual Grounding', Available at: https://omanscience.com/en/articles/m3sunagent-monocular-3d-spatial-understanding-agent-for-metric-depth-estimation-and-3d-visual-grounding.

Vancouver

Zhang J, Wu K, Zhu M, Qiao R, Cai C, Li Z. M3SunAgent: Monocular 3D Spatial Understanding Agent for Metric Depth Estimation and 3D Visual Grounding. https://omanscience.com/en/articles/m3sunagent-monocular-3d-spatial-understanding-agent-for-metric-depth-estimation-and-3d-visual-grounding

IEEE

J. Zhang, K. Wu, M. Zhu, R. Qiao, C. Cai, and Z. Li, "M3SunAgent: Monocular 3D Spatial Understanding Agent for Metric Depth Estimation and 3D Visual Grounding," https://omanscience.com/en/articles/m3sunagent-monocular-3d-spatial-understanding-agent-for-metric-depth-estimation-and-3d-visual-grounding.