نسخة أولية وصول مفتوح
A Bird's-Eye View of Iterative Reward Design
Designing effective reward functions in RL typically requires substantial expertise and trial and error. Recent work automates this process with LLM-based systems that generate and iteratively improve reward code using policy feedback. However, these methods are often hard to compare because they differ in implementati …