الباحثون

Anqi Pu

المنشورات 1

نسخة أولية وصول مفتوح

Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA

Kailong Fan, Anqi Pu, Yichen Wu وآخرون · 2026

Test-time reinforcement learning adapts a model on its own unlabeled test set using majority-vote pseudo-labels and has shown strong results in mathematics. We show that this recipe collapses on medical multiple-choice QA: accuracy stagnates while output diversity rapidly declines. Through a controlled experiment that …

المؤلفون المشاركون