نسخة أولية وصول مفتوح
KVE-KD: Key Visual Evidence-Guided Knowledge Distillation for Vision-Language Models
Knowledge distillation is crucial for deploying vision-language models on resource-constrained devices. However, existing methods typically impose uniform supervision across visual tokens or rely on static token selection, which confuses task-relevant cues with background noise and degrades cross-modal reasoning. To ad …