الباحثون

Qiaosheng Zhang

المنشورات 2

نسخة أولية وصول مفتوح

Semifactual Credit-Augmented Policy Optimization

Junshu Pan, Zhizhang Fu, Shulin Huang وآخرون · 2026

Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its …

المؤلفون المشاركون