الباحثون

Fangxiang Feng

المنشورات 1

نسخة أولية وصول مفتوح

What Makes an Efficient VLA? Navigating Action-Head Design, Scaling, and Latency

Luoyang Sun, Guoyang Xia, Fengfa Li وآخرون · 2026

Vision-Language-Action (VLA) models combine a pretrained vision encoder, a language backbone, and an action head, but their relative contribution has not been established under controlled, latency-paired conditions. We fix the backbone families (SigLIP2 and Qwen2.5) and the training pipeline, sweep action-head design a …

المؤلفون المشاركون