Abstract
Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, this assumption does not hold when cross-modal features are structurally heterogeneous and lack clear unit-level correspondence, such as 2D spatial visual grids and 1D temporal audio sequences, thereby limiting the applicability of feature-level alignment. To address this challenge, we propose a cross-modal distillation framework that enables effective knowledge transfer across structurally heterogeneous feature spaces via a vector-quantized codebook. Specifically, teacher features are abstracted into a set of vector-form codes regardless of their original feature structure, and the selected codes serve as concept-level anchors for student learning. Code selection is guided by both task relevance and student compatibility, allowing the student to receive transferable teacher knowledge without requiring direct unit-level feature alignment. Experimental results across diverse cross-modal distillation scenarios demonstrate the effectiveness of the proposed framework on classification and semantic segmentation tasks.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Jo, D. U., Lim, J., Yoo, Y., & Um, D. (2026). Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features. https://omanscience.com/en/articles/codebook-guided-cross-modal-knowledge-distillation-for-structurally-heterogeneous-features
MLA 9
Jo, Dae Ung, et al. "Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features." https://omanscience.com/en/articles/codebook-guided-cross-modal-knowledge-distillation-for-structurally-heterogeneous-features.
Chicago (author–date)
Jo, Dae Ung, Jongin Lim, YoungJoon Yoo, and Daeho Um. 2026. "Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features." https://omanscience.com/en/articles/codebook-guided-cross-modal-knowledge-distillation-for-structurally-heterogeneous-features.
Harvard
Jo, D. U., Lim, J., Yoo, Y. and Um, D. (2026) 'Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features', Available at: https://omanscience.com/en/articles/codebook-guided-cross-modal-knowledge-distillation-for-structurally-heterogeneous-features.
Vancouver
Jo DU, Lim J, Yoo Y, Um D. Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features. https://omanscience.com/en/articles/codebook-guided-cross-modal-knowledge-distillation-for-structurally-heterogeneous-features
IEEE
D. U. Jo, J. Lim, Y. Yoo, and D. Um, "Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features," https://omanscience.com/en/articles/codebook-guided-cross-modal-knowledge-distillation-for-structurally-heterogeneous-features.