الباحثون

Wentao Ma

المنشورات 3

نسخة أولية وصول مفتوح

Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics

Maoqi Liu, Junwei He, Bowen Zhang وآخرون · 2026

Rubric-based reinforcement learning (Rubric-RL) trains language models where no verifier exists. A judge checks each criterion of a rubric, and the verdicts are aggregated into a reward, most often by a weighted sum. We show that this additive aggregation is the weak point. Under a sum, criteria compensate for one anot …

نسخة أولية وصول مفتوح

T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning

Bowen Zhang, Junwei He, Maoqi Liu وآخرون · 2026

Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajecto …

نسخة أولية وصول مفتوح

MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents

Chenxu Xiong, Dongming Shen, Yuzhi Tang وآخرون · 2026

Voice provides a natural and immediate interface for AI agents. Many settings in which voice agents could be useful, including meetings, households, and collaborative work, are inherently multi-speaker. Supporting these settings introduces challenges that are largely absent from one-on-one interaction. We introduce the …

المؤلفون المشاركون