Abstract

Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, this assumption does not hold when cross-modal features are structurally heterogeneous and lack clear unit-level correspondence, such as 2D spatial visual grids and 1D temporal audio sequences, thereby limiting the applicability of feature-level alignment. To address this challenge, we propose a cross-modal distillation framework that enables effective knowledge transfer across structurally heterogeneous feature spaces via a vector-quantized codebook. Specifically, teacher features are abstracted into a set of vector-form codes regardless of their original feature structure, and the selected codes serve as concept-level anchors for student learning. Code selection is guided by both task relevance and student compatibility, allowing the student to receive transferable teacher knowledge without requiring direct unit-level feature alignment. Experimental results across diverse cross-modal distillation scenarios demonstrate the effectiveness of the proposed framework on classification and semantic segmentation tasks.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Jo, D. U., Lim, J., Yoo, Y., & Um, D. (2026). Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features. https://omanscience.com/en/articles/codebook-guided-cross-modal-knowledge-distillation-for-structurally-heterogeneous-features

MLA 9

Jo, Dae Ung, et al. "Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features." https://omanscience.com/en/articles/codebook-guided-cross-modal-knowledge-distillation-for-structurally-heterogeneous-features.

Chicago (author–date)

Jo, Dae Ung, Jongin Lim, YoungJoon Yoo, and Daeho Um. 2026. "Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features." https://omanscience.com/en/articles/codebook-guided-cross-modal-knowledge-distillation-for-structurally-heterogeneous-features.

Harvard

Jo, D. U., Lim, J., Yoo, Y. and Um, D. (2026) 'Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features', Available at: https://omanscience.com/en/articles/codebook-guided-cross-modal-knowledge-distillation-for-structurally-heterogeneous-features.

Vancouver

Jo DU, Lim J, Yoo Y, Um D. Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features. https://omanscience.com/en/articles/codebook-guided-cross-modal-knowledge-distillation-for-structurally-heterogeneous-features

IEEE

D. U. Jo, J. Lim, Y. Yoo, and D. Um, "Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features," https://omanscience.com/en/articles/codebook-guided-cross-modal-knowledge-distillation-for-structurally-heterogeneous-features.