نسخة أولية وصول مفتوح
CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models
Cinematography, the craft of visual storytelling through framing, lighting, and camera operation, fundamentally shapes how audiences perceive and emotionally engage with video content. While Large Vision Language Models (LVLMs) have made remarkable progress in video question answering, existing benchmarks primarily foc …