Abstract
We present VGGT-Diff, a geometry-routed multi-view diffusion model for sparse-view novel view synthesis. Existing novel view synthesis (NVS) methods face a fundamental trade-off: reconstruction-based approaches preserve observed geometry but struggle to synthesize unseen regions, while diffusion-based methods provide strong generative priors yet rely on implicit source-to-query correspondence. VGGT-Diff bridges these regimes by routing visual geometry latents from VGGT-Ω into a pretrained video diffusion model. Each visual token is associated with a 3D point and confidence, then transformed into query-aligned latent conditions through a confidence-aware Visual Geometry Router (VGR) that preserves front and back surface evidence. These conditions guide joint target-view denoising, while Point-Track Residual Consistency (PTRC) regularizes predicted-clean residuals along reliable 3D tracks, improving multi-view stability. We further introduce robust geometry conditioning, combining training-time regularization with inference-time guidance for improved robustness. Experiments show competitive or state-of-the-art performance across interpolation and extrapolation under different viewpoint difficulties. Our code is available at https://github.com/chenkangjie1123/VGGT-Diff.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Chen, K., Li, X., Zhang, D., Zheng, C., Chen, S., Deng, J., Lin, H., Wai, C. S., Wang, M., Yang, M., Zhong, D., Song, G., Zhang, Y., Liu, X., & Wang, B. (2026). VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis. https://omanscience.com/en/articles/vggt-diff-visual-geometry-meets-diffusion-for-sparse-view-novel-view-synthesis
MLA 9
Chen, Kangjie, et al. "VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis." https://omanscience.com/en/articles/vggt-diff-visual-geometry-meets-diffusion-for-sparse-view-novel-view-synthesis.
Chicago (author–date)
Chen, Kangjie, Xiangyu Li, Dongbin Zhang, Chaoda Zheng, Shijia Chen, Jinhao Deng, Hongbin Lin, Choo Sin Wai, Minqi Wang, Minghao Yang, Dake Zhong, Guorui Song, Yu Zhang, Xianming Liu, and Boyang Wang. 2026. "VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis." https://omanscience.com/en/articles/vggt-diff-visual-geometry-meets-diffusion-for-sparse-view-novel-view-synthesis.
Harvard
Chen, K., Li, X., Zhang, D., Zheng, C., Chen, S., Deng, J., Lin, H., Wai, C. S., Wang, M., Yang, M., Zhong, D., Song, G., Zhang, Y., Liu, X. and Wang, B. (2026) 'VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis', Available at: https://omanscience.com/en/articles/vggt-diff-visual-geometry-meets-diffusion-for-sparse-view-novel-view-synthesis.
Vancouver
Chen K, Li X, Zhang D, Zheng C, Chen S, Deng J, et al. VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis. https://omanscience.com/en/articles/vggt-diff-visual-geometry-meets-diffusion-for-sparse-view-novel-view-synthesis
IEEE
K. Chen, X. Li, D. Zhang, C. Zheng, S. Chen, J. Deng, H. Lin, C. S. Wai, M. Wang, M. Yang, D. Zhong, G. Song, Y. Zhang, X. Liu, and B. Wang, "VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis," https://omanscience.com/en/articles/vggt-diff-visual-geometry-meets-diffusion-for-sparse-view-novel-view-synthesis.