الباحثون

Ran He

المنشورات 6

نسخة أولية وصول مفتوح

From Suppression to Repair: Mitigating Object Hallucination in Large Vision-Language Models via Localized Distribution Alignment

Chen Zhao, Xingping Dong, Jiachun Shi وآخرون · 2026

Object hallucination remains a major obstacle for large vision-language models (LVLMs) to generate reliable content. An intuitive mitigation strategy is to suppress hallucination-related components in hidden representations. However, these components may also contain useful information, and suppressing them can weaken …

نسخة أولية وصول مفتوح

Smaller Models, Better Rejects: Preference Distillation Scaling

Rui Cai, Wenhui Zhu, Xiwen Chen وآخرون · 2026

Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither a …

نسخة أولية وصول مفتوح

GaugeVLM: Structuring Spatial Supervision with Measured Geometric Interventions

Hongbo Wang, Zihan Lin, Wenkui Yang وآخرون · 2026

Vision-language models (VLMs) can contradict themselves across views of the same spatial relation and fail to respond when that relation changes. Addressing these failures requires supervision that captures error magnitude and geometric dependencies across observations, both of which remain implicit in training on indi …

نسخة أولية وصول مفتوح

TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs

Zihan Lin, Xiaohan Wang, Jie Cao وآخرون · 2026

Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or leave the benefit of ex …

نسخة أولية وصول مفتوح

MM-VeriAgent: Learning to Use Extensive Tools to Verify Multimodal Misinformation with Reinforcement Learning

Peipei Li, Shuhan Xia, Shengyang Liu وآخرون · 2026

Real-world multimodal misinformation often involves mixed forgery sources, requiring sample-specific detection strategies. Existing tool-augmented methods rely on predefined workflows or inference-time planning, limiting adaptability or increasing inference cost. To address this issue, we introduce \textbf{MM-VeriAgent …

المؤلفون المشاركون