Preprint Open access
CURIO: Curiosity-Driven Test-Time Learning for Open-Ended Discovery
Open-ended discovery requires learning from repeated attempts while continuing to explore directions whose value is not yet apparent. Search with a frozen large language model (LLM) can reuse previous solutions in context, but cannot update the model from its successes and failures on the test problem. Reinforcement le …