الباحثون

Kai Ye

المنشورات 2

نسخة أولية وصول مفتوح

Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance

Yanyan Zhang, Disheng Liu, Xinpeng Li وآخرون · 2026

While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including manipulated objects, destinations, and backgrounds, is limited by the lack of diversity in robotic training data. Trained end-to-end on such data, VLAs tend to exploit visua …

نسخة أولية وصول مفتوح

From Feed-Forward to Flow: Unifying Reconstruction and Generation Is Easier Than You Think

Haoru Wang, Qianfan Shen, Kai Ye وآخرون · 2026

Reconstruct where the images provide evidence, and generate where they do not: recent success of spatial world models such as Atlas (World Labs Team, 2026) highlights the value of unifying reconstruction and generation in one model. Yet the two have long lived in separate paradigms with distinctive failure modes: feed- …

المؤلفون المشاركون