الباحثون

Pan Lu

المنشورات 7

نسخة أولية وصول مفتوح

Thinking Inertia: LLMs Keep Thinking When Told Not To

Dianqiao Lei, Kevin Qinghong Lin, Pan Lu وآخرون · 2026

Large Language Models (LLMs) increasingly ship with explicit "thinking modes", yet their counterpart, "no-thinking", has received far less attention. We study LLMs' no-thinking behavior along two axes. a. How to measure no-thinking? Prior work typically defines no-thinking through proxies such as a disabled thinking mo …

نسخة أولية وصول مفتوح

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful informat …

نسخة أولية وصول مفتوح

CURIO: Curiosity-Driven Test-Time Learning for Open-Ended Discovery

Tao Feng, Fangxu Yu, Zijie Lei وآخرون · 2026

Open-ended discovery requires learning from repeated attempts while continuing to explore directions whose value is not yet apparent. Search with a frozen large language model (LLM) can reuse previous solutions in context, but cannot update the model from its successes and failures on the test problem. Reinforcement le …

نسخة أولية وصول مفتوح

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

Shuo Xing, Zilin Dai, Chengyuan Qian وآخرون · 2026

While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLM …

نسخة أولية وصول مفتوح

PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress

Kevin Qinghong Lin, Siyuan Hu, Pan Lu وآخرون · 2026

Autoresearch agents are reshaping the research ecosystem, but they can also let flawed claims enter the literature at scale. Human advisors catch such issues in drafts through careful, traceable feedback, yet advisor-style assessment requires extensive manual effort and does not scale. To shift automated paper assessme …

المؤلفون المشاركون