الباحثون

Jiaxin Mao

المنشورات 2

نسخة أولية وصول مفتوح

Learning to Retrieve via Reinforcement Learning in Embedding Space

Qi Liu, Fengming Liang, Yiqun Chen وآخرون · 2026

Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables …

نسخة أولية وصول مفتوح

RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation

Zixuan Yang, Yiqun Chen, Qi Liu وآخرون · 2026

Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuff …

المؤلفون المشاركون