الباحثون

Ziheng Wang

المنشورات 4

نسخة أولية وصول مفتوح

Seeing and Solving Are Not Enough for Vision-Language Models

Ziheng Wang, Mingxuan Xie, Yilin Liu وآخرون · 2026

Vision-language models (VLMs) answer visual questions by combining visual information extraction with downstream problem solving. We investigate a fundamental question: Does an incorrect answer necessarily reflect a failure in visual extraction or problem solving? A model may succeed at both abilities when tested separ …

نسخة أولية وصول مفتوح

VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge

Deyu Cao, Ryuji Oi, Kosuke Matsushima وآخرون · 2026

Billion-parameter vision-language-action (VLA) policies run either onboard, consuming substantial power, or on remote servers, adding communication latency. To address these drawbacks and better balance latency and onboard energy consumption, we propose VLA-ULAP. It partitions inference across decision times, interleav …

المؤلفون المشاركون