نسخة أولية وصول مفتوح
HeiCo-FOCUS: A Clinically Grounded Dataset for Long-Context Video Understanding
Recent advances in Vision-Language Models (VLMs) have led to rapid progress in video understanding across a wide range of benchmark tasks. However, existing evaluations largely focus on short-term reasoning, failing to assess a critical capability: maintaining cumulative temporal consistency over extended time horizons …