Abstract
Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual information. We show that this explanation is incomplete. In a fixed-context Pass@K analysis, repeated sampling from the same pruned visual representation recovers many examples missed by greedy decoding, indicating that useful visual evidence can remain accessible but be used unreliably. We call this the representation-utilization gap. Motivated by this observation, we introduce SCOPD, a sparse-context on-policy self-distillation framework in which a student generates reasoning trajectories from pruned visual tokens while a privileged full-context teacher supervises the same on-policy prefixes. SCOPD requires no ground-truth responses, architectural changes, or additional inference-time computation. We further introduce SCOPD+, which uses a small visual-budget intervention to identify visually sensitive response positions and selectively distill them. At 10% visual-token retention, the Vanilla model retains 86.37% of its unpruned performance across 13 benchmarks. SCOPD raises this to 90.49%, while SCOPD+ further improves it to 92.43%. Across token budgets, benchmarks, and pruning operators, our results show that efficient reasoning depends not only on which visual information survives pruning, but also on how reliably the model learns to use it.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Jeddi, A., Zhang, E., Gerigk, J., Karaimer, H., Azadani, M. N., Luo, J., Le, M. N., Aminian, G., Buurmeijer, H., Chen, Y., Sigal, L., Gilitschenski, I., Derpanis, K. G., Pavone, M., & Taati, B. (2026). SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models. https://omanscience.com/en/articles/scopd-sparse-context-on-policy-self-distillation-for-efficient-vision-language-models
MLA 9
Jeddi, Ahmadreza, et al. "SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models." https://omanscience.com/en/articles/scopd-sparse-context-on-policy-self-distillation-for-efficient-vision-language-models.
Chicago (author–date)
Jeddi, Ahmadreza, Enming Zhang, Jasper Gerigk, Hakki Karaimer, Mozhgan Nasr Azadani, Jiayun Luo, Minh Ngoc Le, Gholamali Aminian, Hugo Buurmeijer, Yongchao Chen, Leonid Sigal, Igor Gilitschenski, Konstantinos G. Derpanis, Marco Pavone, and Babak Taati. 2026. "SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models." https://omanscience.com/en/articles/scopd-sparse-context-on-policy-self-distillation-for-efficient-vision-language-models.
Harvard
Jeddi, A., Zhang, E., Gerigk, J., Karaimer, H., Azadani, M. N., Luo, J., Le, M. N., Aminian, G., Buurmeijer, H., Chen, Y., Sigal, L., Gilitschenski, I., Derpanis, K. G., Pavone, M. and Taati, B. (2026) 'SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models', Available at: https://omanscience.com/en/articles/scopd-sparse-context-on-policy-self-distillation-for-efficient-vision-language-models.
Vancouver
Jeddi A, Zhang E, Gerigk J, Karaimer H, Azadani MN, Luo J, et al. SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models. https://omanscience.com/en/articles/scopd-sparse-context-on-policy-self-distillation-for-efficient-vision-language-models
IEEE
A. Jeddi, E. Zhang, J. Gerigk, H. Karaimer, M. N. Azadani, J. Luo, M. N. Le, G. Aminian, H. Buurmeijer, Y. Chen, L. Sigal, I. Gilitschenski, K. G. Derpanis, M. Pavone, and B. Taati, "SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models," https://omanscience.com/en/articles/scopd-sparse-context-on-policy-self-distillation-for-efficient-vision-language-models.