Abstract

Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual information. We show that this explanation is incomplete. In a fixed-context Pass@K analysis, repeated sampling from the same pruned visual representation recovers many examples missed by greedy decoding, indicating that useful visual evidence can remain accessible but be used unreliably. We call this the representation-utilization gap. Motivated by this observation, we introduce SCOPD, a sparse-context on-policy self-distillation framework in which a student generates reasoning trajectories from pruned visual tokens while a privileged full-context teacher supervises the same on-policy prefixes. SCOPD requires no ground-truth responses, architectural changes, or additional inference-time computation. We further introduce SCOPD+, which uses a small visual-budget intervention to identify visually sensitive response positions and selectively distill them. At 10% visual-token retention, the Vanilla model retains 86.37% of its unpruned performance across 13 benchmarks. SCOPD raises this to 90.49%, while SCOPD+ further improves it to 92.43%. Across token budgets, benchmarks, and pruning operators, our results show that efficient reasoning depends not only on which visual information survives pruning, but also on how reliably the model learns to use it.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Jeddi, A., Zhang, E., Gerigk, J., Karaimer, H., Azadani, M. N., Luo, J., Le, M. N., Aminian, G., Buurmeijer, H., Chen, Y., Sigal, L., Gilitschenski, I., Derpanis, K. G., Pavone, M., & Taati, B. (2026). SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models. https://omanscience.com/en/articles/scopd-sparse-context-on-policy-self-distillation-for-efficient-vision-language-models

MLA 9

Jeddi, Ahmadreza, et al. "SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models." https://omanscience.com/en/articles/scopd-sparse-context-on-policy-self-distillation-for-efficient-vision-language-models.

Chicago (author–date)

Jeddi, Ahmadreza, Enming Zhang, Jasper Gerigk, Hakki Karaimer, Mozhgan Nasr Azadani, Jiayun Luo, Minh Ngoc Le, Gholamali Aminian, Hugo Buurmeijer, Yongchao Chen, Leonid Sigal, Igor Gilitschenski, Konstantinos G. Derpanis, Marco Pavone, and Babak Taati. 2026. "SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models." https://omanscience.com/en/articles/scopd-sparse-context-on-policy-self-distillation-for-efficient-vision-language-models.

Harvard

Jeddi, A., Zhang, E., Gerigk, J., Karaimer, H., Azadani, M. N., Luo, J., Le, M. N., Aminian, G., Buurmeijer, H., Chen, Y., Sigal, L., Gilitschenski, I., Derpanis, K. G., Pavone, M. and Taati, B. (2026) 'SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models', Available at: https://omanscience.com/en/articles/scopd-sparse-context-on-policy-self-distillation-for-efficient-vision-language-models.

Vancouver

Jeddi A, Zhang E, Gerigk J, Karaimer H, Azadani MN, Luo J, et al. SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models. https://omanscience.com/en/articles/scopd-sparse-context-on-policy-self-distillation-for-efficient-vision-language-models

IEEE

A. Jeddi, E. Zhang, J. Gerigk, H. Karaimer, M. N. Azadani, J. Luo, M. N. Le, G. Aminian, H. Buurmeijer, Y. Chen, L. Sigal, I. Gilitschenski, K. G. Derpanis, M. Pavone, and B. Taati, "SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models," https://omanscience.com/en/articles/scopd-sparse-context-on-policy-self-distillation-for-efficient-vision-language-models.