الباحثون

Sung Ju Hwang

المنشورات 7

نسخة أولية وصول مفتوح

AutoSciBench: Autonomous Benchmark Generation for Evaluating Scientific Agents

Dongki Kim, Namkyeong Lee, Surag Nair وآخرون · 2026

As agents rapidly evolve, existing benchmarks can become saturated, limiting their ability to distinguish capabilities and reveal remaining failure modes. Particularly in scientific domains, constructing and updating benchmarks requires substantial time, labor, and domain expertise, making it difficult to keep evaluati …

نسخة أولية وصول مفتوح

Follow the Entities: A Corpus Map for Agentic Search

Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively searching the full cor …

نسخة أولية وصول مفتوح

Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge

Chanuk Lee, Minki Kang, Sangwoo Park وآخرون · 2026

Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple …

نسخة أولية وصول مفتوح

Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning

Woongyeong Yeo, Minki Kang, Chanuk Lee وآخرون · 2026

Reinforcement learning with verifiable rewards (RLVR) enhances reasoning in large language models (LLMs) through outcome-level feedback, yet recent approaches to finer-grained credit assignment often require auxiliary models, additional sampling, or privileged information. Although policy entropy provides a readily ava …

نسخة أولية وصول مفتوح

EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents

Sehee Kim, Yumin Choi, Minki Kang وآخرون · 2026

Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ability to adapt how they gather evidence, invoke tools, verify signals, and manage risk u …

المؤلفون المشاركون