الملخص
Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off: shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion therefore couples two stages that benefit from different information. We introduce FuseReg, which replaces heuristic feature selection with training over random subsets of encoder layers. We theoretically analyze the underlying mechanism: subset sampling preserves the full-layer latent mean in expectation while explicitly penalizing sensitivity to cross-layer disagreement. On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR than decoders specialized to fixed fusions. This flexibility also benefits generation: decoder replacement alone reduces unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator. Applying FuseReg to both stages also reduces unguided gFID by 29% on DiT-Base. The reconstruction and generation benefits also extend to other encoder families. FuseReg narrows the reconstruction-generation gap without additional training cost or architectural changes.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Du, H., Xie, Y., Ye, J., Yang, J., Cong, X., Zhang, H., Huang, Y., Wu, H., Li, Z., Gui, S., Liu, D., Li, R., Ni, J., Wei, C., Balestriero, R., & Wang, Y. (2026). FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders. https://omanscience.com/ar/articles/fusereg-regularizing-layer-fusion-mitigates-the-reconstruction-generation-gap-in-representation-autoencoders
MLA 9
Du, Hongyang, et al. "FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders." https://omanscience.com/ar/articles/fusereg-regularizing-layer-fusion-mitigates-the-reconstruction-generation-gap-in-representation-autoencoders.
شيكاغو (المؤلف–التاريخ)
Du, Hongyang, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong, Haodong Zhang, Yongchao Huang, Haiyu Wu, Zongxia Li, Shihang Gui, Dawei Liu, Runhao Li, Jingcheng Ni, Chen Wei, Randall Balestriero, and Yue Wang. 2026. "FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders." https://omanscience.com/ar/articles/fusereg-regularizing-layer-fusion-mitigates-the-reconstruction-generation-gap-in-representation-autoencoders.
هارفارد
Du, H., Xie, Y., Ye, J., Yang, J., Cong, X., Zhang, H., Huang, Y., Wu, H., Li, Z., Gui, S., Liu, D., Li, R., Ni, J., Wei, C., Balestriero, R. and Wang, Y. (2026) 'FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders', Available at: https://omanscience.com/ar/articles/fusereg-regularizing-layer-fusion-mitigates-the-reconstruction-generation-gap-in-representation-autoencoders.
فانكوفر
Du H, Xie Y, Ye J, Yang J, Cong X, Zhang H, et al. FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders. https://omanscience.com/ar/articles/fusereg-regularizing-layer-fusion-mitigates-the-reconstruction-generation-gap-in-representation-autoencoders
IEEE
H. Du, Y. Xie, J. Ye, J. Yang, X. Cong, H. Zhang, Y. Huang, H. Wu, Z. Li, S. Gui, D. Liu, R. Li, J. Ni, C. Wei, R. Balestriero, and Y. Wang, "FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders," https://omanscience.com/ar/articles/fusereg-regularizing-layer-fusion-mitigates-the-reconstruction-generation-gap-in-representation-autoencoders.