نسخة أولية وصول مفتوح
Do LLMs Learn from Rewards in Context?: Rethinking the role of reward in In-Context Reinforcement Learning
LLM agents increasingly improve at inference time by accumulating experience in context rather than by updating parameters. This process is often described as in-context reinforcement learning (ICRL). Whether in-context learning (ICL) can actually play the role of RL, however, has not been tested. We study this questio …