الباحثون

Zhihan Zhang

المنشورات 4

نسخة أولية وصول مفتوح

Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?

Xingwu Zhang, Duanyang Du, Huiling Zhu وآخرون · 2026

A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU). We show that in-domain training does not close this gap. On MMAD, a widely adopted MM-IAU benchmark, trained specialists reach …

نسخة أولية وصول مفتوح

Rufus-Air: An Open LLM Post-Training Recipe

Chia-Yuan Chang, Renyuan Cheng, Rui Feng وآخرون · 2026

Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and s …

نسخة أولية وصول مفتوح

Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows

Xiyuan Shen, Jiuyang Lyu, Seokhyun Hwang وآخرون · 2026

Video provides a rich record of human behavior, interaction, and situated contexts, offering important evidence for understanding people and conducting human-centered research. As vision-language models (VLMs) become increasingly capable of analyzing video, they offer opportunities to automate this traditionally human- …

المؤلفون المشاركون