الباحثون

Hisham Cholakkal

المنشورات 4

نسخة أولية وصول مفتوح

WorldGuide: Goal-Directed Video World Model for Procedural Task Execution

Ankan Deria, Komal Kumar, Hisham Cholakkal وآخرون · 2026

Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete. Ope …

نسخة أولية وصول مفتوح

Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation

Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosi …

نسخة أولية وصول مفتوح

Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision

Frontier general-purpose systems are rapidly expanding beyond visual understanding into capabilities traditionally handled by dedicated computer-vision models. As these capabilities expand, a central question for the computer-vision community is how far this reach extends, and what remains hard. We evaluate GPT-6 Astra …

المؤلفون المشاركون