الباحثون

Yicong Li

المنشورات 4

نسخة أولية وصول مفتوح

Rethinking Probability-Based Reinforcement Learning From Posterior Concentration

Shiu-Hong Kao, Yubo Zhao, Zhenyu Tian وآخرون · 2026

Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure …

نسخة أولية وصول مفتوح

Benchmarking Vision-Language Models on Synapse Detection and Proofreading in Connectomics

Yicong Li, Junjie Wang, Leander Lauenburg وآخرون · 2026

We benchmarked vision-language models (VLMs) on the decisions annotators take when inspecting electron microscopy images in connectomics: synapse detection (presence and polarity) and proofreading (split errors and merge errors). For synapse detection, we evaluated 19 open and 2 closed models across various architectur …

نسخة أولية وصول مفتوح

Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA

Kailong Fan, Anqi Pu, Yichen Wu وآخرون · 2026

Test-time reinforcement learning adapts a model on its own unlabeled test set using majority-vote pseudo-labels and has shown strong results in mathematics. We show that this recipe collapses on medical multiple-choice QA: accuracy stagnates while output diversity rapidly declines. Through a controlled experiment that …

المؤلفون المشاركون