الملخص
Video provides a rich record of human behavior, interaction, and situated contexts, offering important evidence for understanding people and conducting human-centered research. As vision-language models (VLMs) become increasingly capable of analyzing video, they offer opportunities to automate this traditionally human-intensive process. Yet a central question remains: when can VLMs analyze human-centered video independently, and when does reliable analysis still require human involvement? To address this question, we first characterize video analysis practices in human-centered research. We systematically analyze all 1,702 CHI 2026 full papers and identify 125 that annotate videos. Through iterative coding, we derive a five-dimensional taxonomy spanning analytic purpose, viewpoint, phenomenon, reasoning requirement, and annotation authority. Grounded in recurring annotation tasks captured by this taxonomy, we construct a benchmark of 15 representative tasks from open datasets to map the capabilities and limitations of a general-purpose VLM. We examine the division of labor between humans and VLMs by comparing three annotation workflows: VLM alone, human alone, and human verification of VLM outputs. Across tasks, VLM-alone annotation approaches human accuracy on average (HNS = 97.0, where 100 denotes human-alone performance), demonstrating substantial potential to automate human-centered video analysis. Human verification achieves the highest accuracy (HNS = 121.5) while reducing human annotation time by 48.9% and monetary cost by 31.3%-44.5% relative to human-alone annotation. Our findings connect real-world human-centered video analysis tasks and current VLM capabilities, and clarify how human-AI collaboration can make VLM-assisted analysis reliable and efficient.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Shen, X., Lyu, J., Hwang, S., Yao, H., Patel, S., Zhang, Z., & Wobbrock, J. O. (2026). Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows. https://omanscience.com/ar/articles/can-vision-language-models-analyze-human-centered-video-mapping-model-capabilities-and-human-ai-collaborative-workflows
MLA 9
Shen, Xiyuan, et al. "Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows." https://omanscience.com/ar/articles/can-vision-language-models-analyze-human-centered-video-mapping-model-capabilities-and-human-ai-collaborative-workflows.
شيكاغو (المؤلف–التاريخ)
Shen, Xiyuan, Jiuyang Lyu, Seokhyun Hwang, Huanfen Yao, Shwetak Patel, Zhihan Zhang, and Jacob O. Wobbrock. 2026. "Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows." https://omanscience.com/ar/articles/can-vision-language-models-analyze-human-centered-video-mapping-model-capabilities-and-human-ai-collaborative-workflows.
هارفارد
Shen, X., Lyu, J., Hwang, S., Yao, H., Patel, S., Zhang, Z. and Wobbrock, J. O. (2026) 'Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows', Available at: https://omanscience.com/ar/articles/can-vision-language-models-analyze-human-centered-video-mapping-model-capabilities-and-human-ai-collaborative-workflows.
فانكوفر
Shen X, Lyu J, Hwang S, Yao H, Patel S, Zhang Z, et al. Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows. https://omanscience.com/ar/articles/can-vision-language-models-analyze-human-centered-video-mapping-model-capabilities-and-human-ai-collaborative-workflows
IEEE
X. Shen, J. Lyu, S. Hwang, H. Yao, S. Patel, Z. Zhang, and J. O. Wobbrock, "Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows," https://omanscience.com/ar/articles/can-vision-language-models-analyze-human-centered-video-mapping-model-capabilities-and-human-ai-collaborative-workflows.