Abstract

Digital zoom transitions between dual cameras often exhibit conspicuous discontinuities in geometric structure and chromatic consistency, degrading the user experience. While recent dual-camera smooth zoom (DCSZ) methods attempt to mitigate this by fine-tuning frame interpolation (FI) models on DCSZ data, they struggle with the large cross-view disparities and complex geometric transformations. Considering that the generative prior of diffusion models is suitable for addressing this problem, we explore their application to DCSZ. However, naively applying existing diffusion-based FI models still yields low-fidelity transitions due to insufficient conditional guidance, high-frequency information loss during VAE encoding, as well as inadequate temporal consistency. To address this, we propose ZoomDiff, a high-fidelity diffusion model that leverages dual-camera inputs in both latent and pixel spaces for photo-realistic transitions. Specifically, we first strengthen dual-image conditional guidance during the multi-step denoising process to improve geometric consistency. Then we inject flow-aligned multi-scale features from the VAE encoder into the VAE decoder to recover high-frequency details, where flow-guided temporal consistency supervision are introduced to produce more smooth transitions. Extensive experiments on both synthetic and real-world datasets demonstrate that ZoomDiff outperforms state-of-the-art methods quantitatively and qualitatively. Project page: https://jiayi-hit.github.io/ZoomDiff.github.io/.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhang, J., Wu, R., Ding, Y., Deng, S., & Zuo, W. (2026). ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming. https://omanscience.com/en/articles/zoomdiff-a-high-fidelity-diffusion-model-for-dual-camera-smooth-zooming

MLA 9

Zhang, Jiayi, et al. "ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming." https://omanscience.com/en/articles/zoomdiff-a-high-fidelity-diffusion-model-for-dual-camera-smooth-zooming.

Chicago (author–date)

Zhang, Jiayi, Renlong Wu, Yukang Ding, Sibin Deng, and Wangmeng Zuo. 2026. "ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming." https://omanscience.com/en/articles/zoomdiff-a-high-fidelity-diffusion-model-for-dual-camera-smooth-zooming.

Harvard

Zhang, J., Wu, R., Ding, Y., Deng, S. and Zuo, W. (2026) 'ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming', Available at: https://omanscience.com/en/articles/zoomdiff-a-high-fidelity-diffusion-model-for-dual-camera-smooth-zooming.

Vancouver

Zhang J, Wu R, Ding Y, Deng S, Zuo W. ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming. https://omanscience.com/en/articles/zoomdiff-a-high-fidelity-diffusion-model-for-dual-camera-smooth-zooming

IEEE

J. Zhang, R. Wu, Y. Ding, S. Deng, and W. Zuo, "ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming," https://omanscience.com/en/articles/zoomdiff-a-high-fidelity-diffusion-model-for-dual-camera-smooth-zooming.