الملخص

Sub-billion-parameter Vision-Language Models are increasingly viable for on-device deployment, yet compact model size alone does not guarantee usability. On a phone, such a model can still require more than two minutes to produce its first token. On-device usability depends on three axes: accuracy, efficiency, and behavioral reliability; standard benchmarks miss the third, with answers too short to expose doom loops and prompts too benign to probe adversarial safety. We introduce a diagnosis-driven post-training recipe in which a teacher VLM stress-tests the student, uncovers failure modes beyond human priors, and converts them into targeted supervision and preference alignment, supplementing generic data scaling with failure-driven optimization. Coupled with two visual-token policies, the recipe yields two accuracy-efficiency variants with improved behavioral reliability. \textbf{\NanoFull} attains a 62.3 normalized average over 17 benchmarks, the highest among openly released $\sim$0.5B models, +7.4 over its base at identical architecture and token budget, with doom-loop rates at or below the strongest baseline's. \textbf{\FlashFull} retains 61.4 while cutting warm time-to-first-token on a Pixel 9 from 138\,s to 6.1\,s (23$\times$). By jointly addressing all three axes, we move compact VLMs toward practical on-device usability.

الكلمات المفتاحية

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Hashmi, K. A., Zolfaghari, M., Park, C., Jain, R., Moratelli, N., Wei, P., Lu, L., Liu, T., & Nazir, A. (2026). VisionPsy-Nano: Improving Accuracy, Efficiency, and Reliability in On-Device Vision-Language Models. https://omanscience.com/ar/articles/visionpsy-nano-improving-accuracy-efficiency-and-reliability-in-on-device-vision-language-models

MLA 9

Hashmi, Khurram Azeem, et al. "VisionPsy-Nano: Improving Accuracy, Efficiency, and Reliability in On-Device Vision-Language Models." https://omanscience.com/ar/articles/visionpsy-nano-improving-accuracy-efficiency-and-reliability-in-on-device-vision-language-models.

شيكاغو (المؤلف–التاريخ)

Hashmi, Khurram Azeem, Mohammadreza Zolfaghari, Changdae Park, Rishabh Jain, Nicholas Moratelli, Pengfei Wei, Louis Lu, Tianchi Liu, and Amril Nazir. 2026. "VisionPsy-Nano: Improving Accuracy, Efficiency, and Reliability in On-Device Vision-Language Models." https://omanscience.com/ar/articles/visionpsy-nano-improving-accuracy-efficiency-and-reliability-in-on-device-vision-language-models.

هارفارد

Hashmi, K. A., Zolfaghari, M., Park, C., Jain, R., Moratelli, N., Wei, P., Lu, L., Liu, T. and Nazir, A. (2026) 'VisionPsy-Nano: Improving Accuracy, Efficiency, and Reliability in On-Device Vision-Language Models', Available at: https://omanscience.com/ar/articles/visionpsy-nano-improving-accuracy-efficiency-and-reliability-in-on-device-vision-language-models.

فانكوفر

Hashmi KA, Zolfaghari M, Park C, Jain R, Moratelli N, Wei P, et al. VisionPsy-Nano: Improving Accuracy, Efficiency, and Reliability in On-Device Vision-Language Models. https://omanscience.com/ar/articles/visionpsy-nano-improving-accuracy-efficiency-and-reliability-in-on-device-vision-language-models

IEEE

K. A. Hashmi, M. Zolfaghari, C. Park, R. Jain, N. Moratelli, P. Wei, L. Lu, T. Liu, and A. Nazir, "VisionPsy-Nano: Improving Accuracy, Efficiency, and Reliability in On-Device Vision-Language Models," https://omanscience.com/ar/articles/visionpsy-nano-improving-accuracy-efficiency-and-reliability-in-on-device-vision-language-models.