الملخص

Vision-language foundation models such as CLIP provide strong semantic representations, but their patch tokens are not directly optimized for dense metric geometry. SPACE-CLIP showed that frozen CLIP features can support monocular depth estimation through layer-group feature fusion, yet it leaves open how neighboring CLIP tokens should be combined to recover fine local structure. We present SPACE-CLIPv2, a frozen-backbone depth decoder that aggregates fixed local neighborhoods in CLIP token space. At selected decoder stages, the model samples a fixed token stencil, predicts aggregation weights, and injects the resulting response through a gated residual update. A token-space high-pass branch further preserves shallow local contrast. On NYU Depth V2, SPACE-CLIPv2 improves over a matched SPACE-CLIP baseline, while five-seed experiments consistently favor fixed over learned-offset sampling. Zero-shot iBims-1 evaluation further improves boundary and planar-geometry measures. These results support constrained local token aggregation as a practical mechanism for decoding geometry from frozen CLIP representations.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Song, H., Cho, T., Kim, K., & Choi, A. J. (2026). SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation. https://omanscience.com/ar/articles/space-clipv2-decoding-local-geometry-from-frozen-clip-for-monocular-depth-estimation

MLA 9

Song, Hyun, et al. "SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation." https://omanscience.com/ar/articles/space-clipv2-decoding-local-geometry-from-frozen-clip-for-monocular-depth-estimation.

شيكاغو (المؤلف–التاريخ)

Song, Hyun, Taewan Cho, Kangmin Kim, and Andrew Jaeyong Choi. 2026. "SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation." https://omanscience.com/ar/articles/space-clipv2-decoding-local-geometry-from-frozen-clip-for-monocular-depth-estimation.

هارفارد

Song, H., Cho, T., Kim, K. and Choi, A. J. (2026) 'SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation', Available at: https://omanscience.com/ar/articles/space-clipv2-decoding-local-geometry-from-frozen-clip-for-monocular-depth-estimation.

فانكوفر

Song H, Cho T, Kim K, Choi AJ. SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation. https://omanscience.com/ar/articles/space-clipv2-decoding-local-geometry-from-frozen-clip-for-monocular-depth-estimation

IEEE

H. Song, T. Cho, K. Kim, and A. J. Choi, "SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation," https://omanscience.com/ar/articles/space-clipv2-decoding-local-geometry-from-frozen-clip-for-monocular-depth-estimation.