نسخة أولية وصول مفتوح
The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning
Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. T …