الملخص

Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains underexplored. In this work, we present Flow Any Scene Transformer (FAST), a scalable correspondence model driven by two key insights. First, we reveal that the query-key projections inside single-view vision foundation models encode a coarse yet reusable prior for cross-view matching. Second, reusing these pretrained projections in cross-attention form yields a highly effective initialization for a ViT-based matcher built from a single-view encoder. Guided by these insights, we build FAST upon a vanilla single-view foundation model, utilizing a zero-parameter rewiring strategy to convert selected self-attention layers into cross-attention for cross-view interaction. This design allows ViT-based matchers to scale with advances in single-view foundation models, bypassing the need for a dedicated pair-centric pretraining stage. To fully unlock the scaling potential of this formulation, we assemble a 6-million-pair training corpus for general-purpose dense 2D displacement estimation across diverse co-visible image pairs. Extensive experiments demonstrate that FAST achieves state-of-the-art performance across a wide range of benchmarks, while scaling favorably with both backbone size and training data.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Zhang, Y., Wang, L., Song, Z., Fu, Z., Lin, L., & Guo, Y. (2026). FAST: Flow Any Scene Transformer. https://omanscience.com/ar/articles/fast-flow-any-scene-transformer

MLA 9

Zhang, Yongjian, et al. "FAST: Flow Any Scene Transformer." https://omanscience.com/ar/articles/fast-flow-any-scene-transformer.

شيكاغو (المؤلف–التاريخ)

Zhang, Yongjian, Longguang Wang, Zhuo Song, Zhiheng Fu, Liang Lin, and Yulan Guo. 2026. "FAST: Flow Any Scene Transformer." https://omanscience.com/ar/articles/fast-flow-any-scene-transformer.

هارفارد

Zhang, Y., Wang, L., Song, Z., Fu, Z., Lin, L. and Guo, Y. (2026) 'FAST: Flow Any Scene Transformer', Available at: https://omanscience.com/ar/articles/fast-flow-any-scene-transformer.

فانكوفر

Zhang Y, Wang L, Song Z, Fu Z, Lin L, Guo Y. FAST: Flow Any Scene Transformer. https://omanscience.com/ar/articles/fast-flow-any-scene-transformer

IEEE

Y. Zhang, L. Wang, Z. Song, Z. Fu, L. Lin, and Y. Guo, "FAST: Flow Any Scene Transformer," https://omanscience.com/ar/articles/fast-flow-any-scene-transformer.