[
    {
        "id": "osp-16031",
        "type": "article-journal",
        "title": "Population Scaling or Data Dilution? Dynamics of Local Topology Evolution in Decentralized Learning",
        "author": [
            {
                "family": "Liang",
                "given": "Yin-Kuan"
            },
            {
                "family": "Gao",
                "given": "Yan"
            },
            {
                "family": "Long",
                "given": "Yang"
            }
        ],
        "URL": "https://omanscience.com/en/articles/population-scaling-or-data-dilution-dynamics-of-local-topology-evolution-in-decentralized-learning",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Scaling decentralized learning changes not only the number of clients $N$, but also the dynamics of information propagation and consensus. We argue that the effect of increasing $N$ cannot be understood in isolation, because data allocation, topology-dependent mixing, and communication capacity may change simultaneously. We study these coupled effects on CIFAR-10 with $N\\in\\{10,50,100,200\\}$, comparing a degree-two Ring, a Static Random graph, and Local-First Heuristic Evolution (LFHE), a locally adaptive topology process based on friend-of-friend discovery. The Ring provides an analytically transparent failure mode: its Metropolis spectral gap decays as $Θ(N^{-2})$, implying progressively slower contraction of model disagreement as the population grows. Experiments show that holding the nominal local dataset size fixed substantially reduces the apparent population penalty observed when a fixed total dataset is divided among more clients. The remaining degradation depends strongly on communication structure: Ring enters a high-disagreement regime, whereas Static Random and LFHE remain close to consensus. Increasing LFHE's degree threshold further improves accuracy and consensus, but at a substantially higher model-transmission cost. These results show that decentralized scaling is governed by coupled learning and communication dynamics, rather than by the number of clients alone."
    }
]