الباحثون

Xintong Duan

المنشورات 1

نسخة أولية وصول مفتوح

Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling

Power-sharpened sampling is an inference-time alternative to reinforcement-learning (RL) post-training for enhancing reasoning in large language models (LLMs). High-probability sequences are amplified under the base model without parameter updates or external rewards, avoiding the costly optimization and jagged general …

المؤلفون المشاركون