[
    {
        "id": "osp-19477",
        "type": "article-journal",
        "title": "Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis",
        "author": [
            {
                "family": "Zhang",
                "given": "Wenlong"
            },
            {
                "family": "Jiao",
                "given": "Zhengbo"
            },
            {
                "family": "Zhang",
                "given": "Chenxu"
            },
            {
                "family": "Jiang",
                "given": "Lekang"
            },
            {
                "family": "Ma",
                "given": "Siyuan"
            },
            {
                "family": "Zhang",
                "given": "Qituan"
            },
            {
                "family": "Chen",
                "given": "Guo"
            },
            {
                "family": "Zhang",
                "given": "Linfeng"
            }
        ],
        "URL": "https://omanscience.com/ar/articles/recursive-harness-self-improvement-for-frontier-reasoning-data-synthesis",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Generating progressively harder reasoning problems requires synthesis procedures that adapt as the task distribution evolves. Existing task-level recursion reuses generated problems as seeds but leaves the construction harness unchanged. We present task-harness co-evolution, a framework for recursive harness self-improvement (RSI) in reasoning-data synthesis. Online self-improvement converts intermediate solver failures into reusable skills during generation. Post-task self-improvement revises skills, prompts, and workflows after each batch, adopting candidates only when they generate harder valid tasks within a bounded cost increase. Model weights and verification criteria remain fixed. Across mathematics, coding, and science, mean solver accuracy decreases from 100.0% to 54.8% over fourteen evolution rounds. Ablations show that combining both update schedules produces harder tasks than fixed-harness recursion or either schedule alone. The resulting data improves downstream SFT and GRPO performance. In particular, a 27B student fine-tuned on 10K synthesized mathematics examples achieves 62.5% mean-16 accuracy on APEX, competitive with selected frontier-model references. These results support adapting the synthesis harness alongside the tasks to generate increasingly challenging data with downstream training value."
    }
]