Abstract

Recovering detailed geometry from high-resolution images is critical for precise perception of the surroundings and objects. However, existing methods which use latent-space modeling and VAE reconstruction can compromise geometric details. Furthermore, decoding from latent codes introduces substantial inference overhead. To address those issues, we present EagleDepth, an efficient framework for high-resolution monocular depth estimation that combines the geometric priors of latent diffusion with fine-grained pixel-space generation. Our key idea is to retain depth-aware latent representations as guidance while generating the final depth map directly in pixel space. We train the latent and pixel components sequentially: first, we fine-tune a pretrained latent diffusion model using paired RGB--depth supervision; then, we adapt a pretrained pixel diffusion decoder, PiD, to predict depth conditioned on the learned features. Training of the pixel component starts at 1024 resolution and continues across multiple resolutions up to 4K. The latent branch processes resized, lower-resolution RGB images, while the pixel branch generates depth at the target resolution, bypassing the original VAE decoder. This design preserves learned geometric knowledge without requiring the latent backbone to operate at the output resolution. On five commonly used depth estimation datasets and the high-resolution Synth4K dataset, our framework achieves state-of-the-art depth estimation performance, with faster inference and better preservation of fine structures and object boundaries.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Chai, B., Zhang, T., Wu, S., Zuo, D., Fan, Z., & Zou, D. (2026). EagleDepth: Efficient Fine-Grained Depth Estimation via Pixel Diffusion Decoder. https://omanscience.com/en/articles/eagledepth-efficient-fine-grained-depth-estimation-via-pixel-diffusion-decoder

MLA 9

Chai, Bowen, et al. "EagleDepth: Efficient Fine-Grained Depth Estimation via Pixel Diffusion Decoder." https://omanscience.com/en/articles/eagledepth-efficient-fine-grained-depth-estimation-via-pixel-diffusion-decoder.

Chicago (author–date)

Chai, Bowen, Tianbao Zhang, Shuyu Wu, Dexin Zuo, Zhaoxin Fan, and Danping Zou. 2026. "EagleDepth: Efficient Fine-Grained Depth Estimation via Pixel Diffusion Decoder." https://omanscience.com/en/articles/eagledepth-efficient-fine-grained-depth-estimation-via-pixel-diffusion-decoder.

Harvard

Chai, B., Zhang, T., Wu, S., Zuo, D., Fan, Z. and Zou, D. (2026) 'EagleDepth: Efficient Fine-Grained Depth Estimation via Pixel Diffusion Decoder', Available at: https://omanscience.com/en/articles/eagledepth-efficient-fine-grained-depth-estimation-via-pixel-diffusion-decoder.

Vancouver

Chai B, Zhang T, Wu S, Zuo D, Fan Z, Zou D. EagleDepth: Efficient Fine-Grained Depth Estimation via Pixel Diffusion Decoder. https://omanscience.com/en/articles/eagledepth-efficient-fine-grained-depth-estimation-via-pixel-diffusion-decoder

IEEE

B. Chai, T. Zhang, S. Wu, D. Zuo, Z. Fan, and D. Zou, "EagleDepth: Efficient Fine-Grained Depth Estimation via Pixel Diffusion Decoder," https://omanscience.com/en/articles/eagledepth-efficient-fine-grained-depth-estimation-via-pixel-diffusion-decoder.