الباحثون

Chen Zhao

المنشورات 8

نسخة أولية وصول مفتوح

From Suppression to Repair: Mitigating Object Hallucination in Large Vision-Language Models via Localized Distribution Alignment

Chen Zhao, Xingping Dong, Jiachun Shi وآخرون · 2026

Object hallucination remains a major obstacle for large vision-language models (LVLMs) to generate reliable content. An intuitive mitigation strategy is to suppress hallucination-related components in hidden representations. However, these components may also contain useful information, and suppressing them can weaken …

نسخة أولية وصول مفتوح

ColdDDI: Evaluating Knowledge Utilization in Cold-Start Drug-Drug Interaction Prediction

Jiheng Liang, Chen Zhao, Di Wu وآخرون · 2026

Cold-start drug-drug interaction (DDI) prediction tests whether models can identify clinically significant interactions for drugs without training-time interaction history. Existing benchmarks mostly report aggregate edge-prediction scores, leaving a key evaluation question unanswered: when models receive molecular, te …

نسخة أولية وصول مفتوح

Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment

Yihuai Hong, Shauli Ravfogel, Chen Zhao وآخرون · 2026

Chain-of-thought (CoT) traces often serve as a proxy for how Large Language Models (LLMs) arrive at their answers. However, growing evidence shows that models' CoT often fails to reflect their internal computations and can be changed without affecting their final answers. In this work, we measure and improve the alignm …

نسخة أولية وصول مفتوح

RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation

Shifeng Bao, Fanding Huang, Yihan Lin وآخرون · 2026

Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop ha …

نسخة أولية وصول مفتوح

PGL-3D: Towards Progressive Geometric Learning for 3D Visual Query Localization

Liang Peng, Shizhuo Mu, Bohan Tan وآخرون · 2026

3D Visual Query Localization (3DVQL) retrieves the latest contiguous occurrence of a queried object in an RGB--point-cloud sequence and predicts a 9-DoF cuboid for every response frame. The query is captured independently of the search sequence, so its annotated pose may differ from how the object appears in the search …

نسخة أولية وصول مفتوح

Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling

Guanlin Li, Shifeng Bao, Yihan Zhao وآخرون · 2026

Achieving robust cross-embodiment generalization in imitation learning demands overcoming a critical representation flaw that inextricably entangles task semantics with hardware-specific visual geometry. We propose an interaction-centric framework that leverages the shared structure of two-finger grippers via a paramet …

المؤلفون المشاركون