الباحثون

Guangtao Zhai

المنشورات 4

نسخة أولية وصول مفتوح

From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs

Zijian Chen, Zhengyu Chen, Bohan Liang وآخرون · 2026

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and code generation. However, existing benchmarks typically evaluate these two modalities in isolation, lacking a dedicated assessment of their unification, i.e., how a model can perceive complex visual struc …

نسخة أولية وصول مفتوح

Harnessing Multimodal Large Language Models for Training-Free Human-Object Interaction Detection

Zhaolin Cai, Huiyu Duan, Liu Yang وآخرون · 2026

Human-object interaction (HOI) detection aims to localize human-object pairs and recognize their interactions. Traditional supervised methods perform strongly but rely on task-specific training. Recent multimodal large language models (MLLMs) offer a promising route to training-free HOI detection through their broad vi …

نسخة أولية وصول مفتوح

Evaluating the Evaluators: Diagnosing Large Multimodal Models for AI-Generated Image Assessment

Yu Zhao, Jiarui Wang, Huiyu Duan وآخرون · 2026

With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to capture fine-grained alignment and generative artifacts. While large multimodal models (LMMs) are increasingly adopted as evaluators, existing benchmarks typically study seman …

نسخة أولية وصول مفتوح

Invisible in Space, Visible in Time: Motion Vision CAPTCHA against GUI Agents

Zeyu Zhang, Dingyi Rong, Zijian Chen وآخرون · 2026

Most existing visual CAPTCHAs remain spatially solvable: the required information is exposed by static appearance, local structure, and interface state. This assumption is weakened by advances in multimodal large language models (MLLMs) and Graphical User Interface (GUI) agents, which exhibit strong visual perception, …

المؤلفون المشاركون