الباحثون

Vipin Chaudhary

المنشورات 3

نسخة أولية وصول مفتوح

Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance

Yanyan Zhang, Disheng Liu, Xinpeng Li وآخرون · 2026

While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including manipulated objects, destinations, and backgrounds, is limited by the lack of diversity in robotic training data. Trained end-to-end on such data, VLAs tend to exploit visua …

نسخة أولية وصول مفتوح

How to Loop MoE: Flatten the Experts, Untie the Attention

Shouren Wang, Chuang Ma, Mohsen Hariri وآخرون · 2026

Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and …

نسخة أولية وصول مفتوح

Toward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language Models

Tuo Liang, Disheng Liu, Nengbo Wang وآخرون · 2026

Grounding is a core capability of spatial vision-language models, yet most existing work focuses only on where a referred object is. Many 3D tasks also require knowing how it is oriented. Although existing 3D VLMs may predict oriented boxes, box pose does not explicitly capture object-centric orientation or symmetry-in …

المؤلفون المشاركون