الباحثون

Omar Attia

المنشورات 2

نسخة أولية وصول مفتوح

RLTL;DR: Self-improvement by Internalizing Self-generated Feedback

The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and wh …

المؤلفون المشاركون