الباحثون

Xia Hu

المنشورات 5

نسخة أولية وصول مفتوح

Before Agent Tells The Lie: Has Deception Already Been Represented?

Xinling Li, Dadi Guo, Qingyu Liu وآخرون · 2026

Large language model (LLM)-based agents can exhibit deceptive behavior during task execution, including hiding failures, fabricating results, or falsely signaling task completion. Existing monitoring approaches mainly detect deception after it appears in observable actions or outputs. In this paper, we investigate whet …

نسخة أولية وصول مفتوح

Do More Modalities Always Help? A Geometric Perspective on Missing-Modality Robustness

Songyuan Sui, Zhen Tan, Mohan Zhang وآخرون · 2026

Missing modality remains a longstanding challenge in multimodal learning. Existing methods typically address this issue through modality recovery or adaptive strategies. However, they overlook models' internal cross-modal dependencies formed during multimodal training, which later impair robustness. We systematically c …

نسخة أولية وصول مفتوح

Where MLLMs Fail and Why: Causal Task Decomposition for Capability Failure Diagnosis

Xia Hu, Brian Potetz, Chun-Ta Lu وآخرون · 2026

End-to-end accuracy on compositional tasks records how often MLLMs fail, but cannot distinguish whether a failure reflects an intrinsic deficit in the targeted capability or a cascading error from an upstream prerequisite. We propose a causal decomposition framework that isolates these two failure modes through control …

نسخة أولية وصول مفتوح

RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback

Jingyi He, Nier Wu, Shuang Liu وآخرون · 2026

Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the response behaviors associated with their scoring decisions. Exi …

المؤلفون المشاركون