الباحثون

Shuhan Zhong

المنشورات 3

نسخة أولية وصول مفتوح

Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics

Maoqi Liu, Junwei He, Bowen Zhang وآخرون · 2026

Rubric-based reinforcement learning (Rubric-RL) trains language models where no verifier exists. A judge checks each criterion of a rubric, and the verdicts are aggregated into a reward, most often by a weighted sum. We show that this additive aggregation is the weak point. Under a sum, criteria compensate for one anot …

نسخة أولية وصول مفتوح

T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning

Bowen Zhang, Junwei He, Maoqi Liu وآخرون · 2026

Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajecto …

نسخة أولية وصول مفتوح

RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications

Jianpeng Zhao, Haihua Xu, Haoyang Zhang وآخرون · 2026

We study reasoning in Rule-Governed Decision Tasks (RGDTs), where models apply external rules to case facts and justify decisions, as required in policy, contract, and compliance settings. Beyond the deductive capability emphasized by standard mathematical and logical reasoning tasks, RGDTs require interpreting rules a …

المؤلفون المشاركون