الباحثون

Salman Khan

المنشورات 8

نسخة أولية وصول مفتوح

WorldGuide: Goal-Directed Video World Model for Procedural Task Execution

Ankan Deria, Komal Kumar, Hisham Cholakkal وآخرون · 2026

Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete. Ope …

نسخة أولية وصول مفتوح

HeiCo-FOCUS: A Clinically Grounded Dataset for Long-Context Video Understanding

Leon Mayer, Lucas Luttner, Patrick Godau وآخرون · 2026

Recent advances in Vision-Language Models (VLMs) have led to rapid progress in video understanding across a wide range of benchmark tasks. However, existing evaluations largely focus on short-term reasoning, failing to assess a critical capability: maintaining cumulative temporal consistency over extended time horizons …

نسخة أولية وصول مفتوح

Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models

Qi Lyu, Jiahua Dong, Hao Shen وآخرون · 2026

World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redu …

نسخة أولية وصول مفتوح

Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision

Frontier general-purpose systems are rapidly expanding beyond visual understanding into capabilities traditionally handled by dedicated computer-vision models. As these capabilities expand, a central question for the computer-vision community is how far this reach extends, and what remains hard. We evaluate GPT-6 Astra …

نسخة أولية وصول مفتوح

Program-Verified Self-Evolution for Vision-Language Models

Ahmed Heakl, Sungik Choi, Moontae Lee وآخرون · 2026

Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24\% of majority-vote labels and 18\% of model-judge labels pro …

نسخة أولية وصول مفتوح

Recipe-Matching, Not Equivalence

MathNet-Retrieve asks a retriever to find, for a math problem, a document stating the same problem. An LLM under one fixed prompt writes each gold document and its near-miss distractors; LLM judges filter them. We call this procedure the "recipe", training on pairs built the same way "recipe-matching", and ask how much …

نسخة أولية وصول مفتوح

Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams

One long-form exam in a large course costs hundreds of grader-hours, and qualified graders are scarce; LLM graders are a tempting alternative. To show its pitfalls we grade a practical Computer Vision exam ($570$ dual-graded students) under $171$ configurations spanning closed and open-weights models; the best reaches …

نسخة أولية وصول مفتوح

The Past Frames the Future: Memory for Autoregressive Video Generation

Harold Haodong Chen, Rongjin Guo, Disen Lan وآخرون · 2026

Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands, …

المؤلفون المشاركون