الملخص
Vision-language foundation models such as CLIP provide strong semantic representations, but their patch tokens are not directly optimized for dense metric geometry. SPACE-CLIP showed that frozen CLIP features can support monocular depth estimation through layer-group feature fusion, yet it leaves open how neighboring CLIP tokens should be combined to recover fine local structure. We present SPACE-CLIPv2, a frozen-backbone depth decoder that aggregates fixed local neighborhoods in CLIP token space. At selected decoder stages, the model samples a fixed token stencil, predicts aggregation weights, and injects the resulting response through a gated residual update. A token-space high-pass branch further preserves shallow local contrast. On NYU Depth V2, SPACE-CLIPv2 improves over a matched SPACE-CLIP baseline, while five-seed experiments consistently favor fixed over learned-offset sampling. Zero-shot iBims-1 evaluation further improves boundary and planar-geometry measures. These results support constrained local token aggregation as a practical mechanism for decoding geometry from frozen CLIP representations.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Song, H., Cho, T., Kim, K., & Choi, A. J. (2026). SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation. https://omanscience.com/ar/articles/space-clipv2-decoding-local-geometry-from-frozen-clip-for-monocular-depth-estimation
MLA 9
Song, Hyun, et al. "SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation." https://omanscience.com/ar/articles/space-clipv2-decoding-local-geometry-from-frozen-clip-for-monocular-depth-estimation.
شيكاغو (المؤلف–التاريخ)
Song, Hyun, Taewan Cho, Kangmin Kim, and Andrew Jaeyong Choi. 2026. "SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation." https://omanscience.com/ar/articles/space-clipv2-decoding-local-geometry-from-frozen-clip-for-monocular-depth-estimation.
هارفارد
Song, H., Cho, T., Kim, K. and Choi, A. J. (2026) 'SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation', Available at: https://omanscience.com/ar/articles/space-clipv2-decoding-local-geometry-from-frozen-clip-for-monocular-depth-estimation.
فانكوفر
Song H, Cho T, Kim K, Choi AJ. SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation. https://omanscience.com/ar/articles/space-clipv2-decoding-local-geometry-from-frozen-clip-for-monocular-depth-estimation
IEEE
H. Song, T. Cho, K. Kim, and A. J. Choi, "SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation," https://omanscience.com/ar/articles/space-clipv2-decoding-local-geometry-from-frozen-clip-for-monocular-depth-estimation.