الباحثون

Tong Wei

المنشورات 4

نسخة أولية وصول مفتوح

Dimension-Free Decentralized Nonsmooth Nonconvex Stochastic Optimization

Yuanyu Wan, Lan Xue, Haomin Bai وآخرون · 2026

We investigate decentralized nonsmooth nonconvex stochastic optimization over a network of $n$ nodes, with the goal of finding an $(δ,ε)$-Goldstein stationary point. The best existing algorithm achieves $O(δ^{-1}(ε^{-3}+dε^{-1}))$ sample complexity and $\widetilde{O}(γ^{-1/2}δ^{-1}(ε^{-3}+dε^{-1}))$ communication compl …

نسخة أولية وصول مفتوح

Outcome-Guided On-Policy Self-Distillation

ZheXu Wang, Mao-Lin Luo, Yankun Hong وآخرون · 2026

On-policy self-distillation (OPSD) provides denser token-level supervision and better computational efficiency than Reinforcement Learning with Verifiable Rewards (RLVR). However, this denser supervision may introduce substantial noise and training instability. Existing improvements often rely on high-variance per-toke …

نسخة أولية وصول مفتوح

Learning Process Rewards via Reasoning State Propagation

Kai Gan, Zi-Hao Zhou, Bo Ye وآخرون · 2026

Process reward models (PRMs) have demonstrated notable effectiveness in test-time scaling and reinforcement learning by providing fine-grained signals for evaluating intermediate reasoning states, but their training relies heavily on costly process annotations. A natural way to alleviate this dependence is to complemen …

المؤلفون المشاركون