الباحثون

Fahad Shahbaz Khan

المنشورات 3

نسخة أولية وصول مفتوح

WorldGuide: Goal-Directed Video World Model for Procedural Task Execution

Ankan Deria, Komal Kumar, Hisham Cholakkal وآخرون · 2026

Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete. Ope …

نسخة أولية وصول مفتوح

Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models

Qi Lyu, Jiahua Dong, Hao Shen وآخرون · 2026

World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redu …

نسخة أولية وصول مفتوح

Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision

Frontier general-purpose systems are rapidly expanding beyond visual understanding into capabilities traditionally handled by dedicated computer-vision models. As these capabilities expand, a central question for the computer-vision community is how far this reach extends, and what remains hard. We evaluate GPT-6 Astra …

المؤلفون المشاركون