الملخص
We present a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. We argue that this is not only a modeling problem but also a representation problem: generators typically evolve appearance-centric latents, while perception models recover geometry in a semantically rich space that encodes cross-view structure. Rather than adding geometry as another output, we reparameterize a geometry foundation model's features into a compact latent space for generation. We realize this shift with the geometry-native autoencoder (GAE), whose latent is jointly decodable to appearance, depth, cameras, and point maps. With this state, a standard conditional flow supports diverse generation tasks. In controlled comparisons that hold the generator and training protocol fixed, replacing the latent with GAE improves both visual quality and independently measured 3D coherence: FVD falls by $12.7\%$ and $23.1\%$ on RealEstate10K and DL3DV, and camera-trajectory error is halved on RealEstate10K. Together, these results show that the latent space is central to geometry-consistent generation and can serve as a shared interface between perception and generation.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Lu, J., Yin, M., Hu, W., Liu, H., Zhao, W., Yeung, S. K., Shan, Y., & Liu, Y. (2026). GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation. https://omanscience.com/ar/articles/gae-learning-a-geometry-native-latent-space-for-3d-consistent-world-generation
MLA 9
Lu, Jiahao, et al. "GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation." https://omanscience.com/ar/articles/gae-learning-a-geometry-native-latent-space-for-3d-consistent-world-generation.
شيكاغو (المؤلف–التاريخ)
Lu, Jiahao, Minghao Yin, Wenbo Hu, Hengyu Liu, Wang Zhao, Sai-Kit Yeung, Ying Shan, and Yuan Liu. 2026. "GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation." https://omanscience.com/ar/articles/gae-learning-a-geometry-native-latent-space-for-3d-consistent-world-generation.
هارفارد
Lu, J., Yin, M., Hu, W., Liu, H., Zhao, W., Yeung, S. K., Shan, Y. and Liu, Y. (2026) 'GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation', Available at: https://omanscience.com/ar/articles/gae-learning-a-geometry-native-latent-space-for-3d-consistent-world-generation.
فانكوفر
Lu J, Yin M, Hu W, Liu H, Zhao W, Yeung SK, et al. GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation. https://omanscience.com/ar/articles/gae-learning-a-geometry-native-latent-space-for-3d-consistent-world-generation
IEEE
J. Lu, M. Yin, W. Hu, H. Liu, W. Zhao, S. K. Yeung, Y. Shan, and Y. Liu, "GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation," https://omanscience.com/ar/articles/gae-learning-a-geometry-native-latent-space-for-3d-consistent-world-generation.