الباحثون

Woongyeong Yeo

المنشورات 2

نسخة أولية وصول مفتوح

Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge

Chanuk Lee, Minki Kang, Sangwoo Park وآخرون · 2026

Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple …

نسخة أولية وصول مفتوح

Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning

Woongyeong Yeo, Minki Kang, Chanuk Lee وآخرون · 2026

Reinforcement learning with verifiable rewards (RLVR) enhances reasoning in large language models (LLMs) through outcome-level feedback, yet recent approaches to finer-grained credit assignment often require auxiliary models, additional sampling, or privileged information. Although policy entropy provides a readily ava …

المؤلفون المشاركون