الباحثون

Chi Bene Chen

المنشورات 2

نسخة أولية وصول مفتوح

VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models

Owen Du, Yang Yue, Jie Zhang وآخرون · 2026

Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but incur high computational costs from processing long token sequences at every control step, limiting real-time deployment. Visual token pruning offers a direct solution, as visual patches dominate the input sequence and contain consi …

نسخة أولية وصول مفتوح

What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling

Renping Zhou, Zanlin Ni, Zihao Fan وآخرون · 2026

World action models (WAMs) predict the future alongside actions during training. Due to the heavy computation cost of video denoising, whether the future must still be generated during inference is disputed: Explicit WAMs denoise it into clean frames along with every action chunk, whereas Latent WAMs discard it entirel …

المؤلفون المشاركون