Abstract

Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains underexplored. In this work, we present Flow Any Scene Transformer (FAST), a scalable correspondence model driven by two key insights. First, we reveal that the query-key projections inside single-view vision foundation models encode a coarse yet reusable prior for cross-view matching. Second, reusing these pretrained projections in cross-attention form yields a highly effective initialization for a ViT-based matcher built from a single-view encoder. Guided by these insights, we build FAST upon a vanilla single-view foundation model, utilizing a zero-parameter rewiring strategy to convert selected self-attention layers into cross-attention for cross-view interaction. This design allows ViT-based matchers to scale with advances in single-view foundation models, bypassing the need for a dedicated pair-centric pretraining stage. To fully unlock the scaling potential of this formulation, we assemble a 6-million-pair training corpus for general-purpose dense 2D displacement estimation across diverse co-visible image pairs. Extensive experiments demonstrate that FAST achieves state-of-the-art performance across a wide range of benchmarks, while scaling favorably with both backbone size and training data.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhang, Y., Wang, L., Song, Z., Fu, Z., Lin, L., & Guo, Y. (2026). FAST: Flow Any Scene Transformer. https://omanscience.com/en/articles/fast-flow-any-scene-transformer

MLA 9

Zhang, Yongjian, et al. "FAST: Flow Any Scene Transformer." https://omanscience.com/en/articles/fast-flow-any-scene-transformer.

Chicago (author–date)

Zhang, Yongjian, Longguang Wang, Zhuo Song, Zhiheng Fu, Liang Lin, and Yulan Guo. 2026. "FAST: Flow Any Scene Transformer." https://omanscience.com/en/articles/fast-flow-any-scene-transformer.

Harvard

Zhang, Y., Wang, L., Song, Z., Fu, Z., Lin, L. and Guo, Y. (2026) 'FAST: Flow Any Scene Transformer', Available at: https://omanscience.com/en/articles/fast-flow-any-scene-transformer.

Vancouver

Zhang Y, Wang L, Song Z, Fu Z, Lin L, Guo Y. FAST: Flow Any Scene Transformer. https://omanscience.com/en/articles/fast-flow-any-scene-transformer

IEEE

Y. Zhang, L. Wang, Z. Song, Z. Fu, L. Lin, and Y. Guo, "FAST: Flow Any Scene Transformer," https://omanscience.com/en/articles/fast-flow-any-scene-transformer.