Authors

Jinwoo Kim

Publications 5

Preprint Open access

Understanding Enrichment in Reinforcement Learning

When rewards are sparse, reinforcement learning with verifiable rewards (RLVR) often uses hints or intermediate guidance to generate more successful rollouts. This enrichment biases policy-gradient updates unless corrected via importance weights, but existing methods omit correction or truncate importance weights in or …

Preprint Open access

Constraint-Aware Training

Jinwoo Kim · 2026

When generating programs with language models, constrained decoding can apply program analyses to exclude tokens that violate syntax, scope, or typing rules. However, there is a duplication: standard training already teaches the model to suppress the tokens rejected by these analyses. This duplication leads to the ques …

Co-authors