[
    {
        "id": "osp-20527",
        "type": "article-journal",
        "title": "RLTL;DR: Self-improvement by Internalizing Self-generated Feedback",
        "author": [
            {
                "family": "Kirchhof",
                "given": "Michael"
            },
            {
                "family": "Gualdoni",
                "given": "Eleonora"
            },
            {
                "family": "Szot",
                "given": "Andrew"
            },
            {
                "family": "Gatmiry",
                "given": "Khashayar"
            },
            {
                "family": "Lotfi",
                "given": "Aryo"
            },
            {
                "family": "Kazerouni",
                "given": "Abbas"
            },
            {
                "family": "Attia",
                "given": "Omar"
            },
            {
                "family": "Chowdhury",
                "given": "Sanjoy"
            },
            {
                "family": "Toshev",
                "given": "Alexander"
            }
        ],
        "URL": "https://omanscience.com/en/articles/rltl-dr-self-improvement-by-internalizing-self-generated-feedback",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned on all previous insights, and we sequentially sample rollouts until a solution is found. Moreover, we enable backpropagation on the in-context insights to internalize a direct task to insight mapping. On challenging tool-calling and coding datasets (filtered to Pass@128=0), standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR breaks through this learning barrier, achieving a Pass@1 of 14-31% with insights in context during training and, crucially, 12-13% when no insight is in context at eval time. We identify that the key is the task to insight internalization. To study this further, we reduce our approach to SFTL;DR, training only on (task, insight) tuples, without showing or backpropagating on any rollouts. Training on only 4k of these tuples recovers almost the full performance of RLTL;DR and classical SFT on full rollouts. This demonstrates a promising compacted training paradigm of the form \"on this sort of task, keep this sort of thing in mind\", which we hope to inspire future research on."
    }
]