[
    {
        "id": "osp-16071",
        "type": "article-journal",
        "title": "Fusion is the New Mutation: Bandit-Guided Evolution on Workflow Graphs",
        "author": [
            {
                "family": "Shang",
                "given": "Zhiwei"
            },
            {
                "family": "Sun",
                "given": "Jiahang"
            },
            {
                "family": "Gong",
                "given": "Mingrong"
            },
            {
                "family": "Kong",
                "given": "Mingze"
            },
            {
                "family": "Qu",
                "given": "Zikun"
            },
            {
                "family": "Lu",
                "given": "Pingchen"
            },
            {
                "family": "Dong",
                "given": "Junhao"
            },
            {
                "family": "Liu",
                "given": "Zhipiao"
            },
            {
                "family": "Yang",
                "given": "Hongwei"
            },
            {
                "family": "Xie",
                "given": "Guoqing"
            },
            {
                "family": "Shu",
                "given": "Yao"
            },
            {
                "family": "Dai",
                "given": "Zhongxiang"
            }
        ],
        "URL": "https://omanscience.com/ar/articles/fusion-is-the-new-mutation-bandit-guided-evolution-on-workflow-graphs",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Automated agentic workflow optimization relies on costly evaluations, making it essential to allocate a limited evaluation budget effectively. Multi-parent fusion can reuse designs from previously discovered workflows, but identifying promising parent combinations requires learning from limited fusion feedback. We introduce DAGO (Directed Acyclic Graph Optimization), a contextual-bandit-guided framework that learns which parent workflows to fuse under a limited evaluation budget. DAGO formulates each candidate parent combination as an arm, represented by pretrained embeddings of its constituent workflows' code and prompts. A diagonal LinUCB policy learns a shared reward model across arms and balances exploitation of arms with high predicted offspring quality against uncertainty-driven exploration. After an arm is selected, an LLM generates a child workflow through summary-guided fusion, and the child's validation score serves as the reward for updating the bandit. A shared directed acyclic graph maintains discovered workflows and their multi-parent lineage, providing an expanding pool of parents for subsequent arm proposals. Across six benchmarks covering mathematical reasoning, code generation, and question answering, DAGO achieves the highest macro-average score among the evaluated baselines. Under matched validation-evaluation budgets, it improves over AFlow from 80.3 to 81.7 while reducing aggregate search expenditure by 11.2%. Ablation studies show that LinUCB-guided arm selection outperforms both random selection and its exploration-free variant, supporting the value of feedback-driven selection and exploration-exploitation balance."
    }
]