نسخة أولية وصول مفتوح
Learning to Plan by Looking Back: Hindsight Hierarchies for Training Reasoning Models
We introduce a self-improvement loop for reasoning models based on the following observation: Even when the difficulty of a problem exceeds the model's current solving abilities, an additionally supplied solution might enable the model to extract useful solution ideas in hindsight. We operationalize this by jointly tra …