الباحثون

Yan Gao

المنشورات 4

نسخة أولية وصول مفتوح

LFHE: Local-First Heuristic Evolution for Bounded Local Topology Search in Decentralized Learning with Non-IID Data

Decentralized learning is highly sensitive to communication topology under non-IID data. Adaptive peer-selection methods can exploit local model information, but broader peer discovery may require increasingly large control state, whereas direct spectral optimization typically relies on graph-wide information. We study …

نسخة أولية وصول مفتوح

GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents

Bohan Lin, Liyi Chen, Zhuoning Guo وآخرون · 2026

On-policy distillation post-trains large language model agents by supplying dense, step-level guidance from a teacher policy when the reinforcement-learning reward is sparse and arrives only once per trajectory. Existing instantiations allocate this guidance by the size of the teacher-student divergence at each step, o …

نسخة أولية وصول مفتوح

Population Scaling or Data Dilution? Dynamics of Local Topology Evolution in Decentralized Learning

Scaling decentralized learning changes not only the number of clients $N$, but also the dynamics of information propagation and consensus. We argue that the effect of increasing $N$ cannot be understood in isolation, because data allocation, topology-dependent mixing, and communication capacity may change simultaneousl …

نسخة أولية وصول مفتوح

RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation

Zixuan Yang, Yiqun Chen, Qi Liu وآخرون · 2026

Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuff …

المؤلفون المشاركون