نسخة أولية وصول مفتوح
Base Models Can Reason By Taking a Cue From Training Data
In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-tr …