الباحثون

Yuhua Zhou

المنشورات 2

نسخة أولية وصول مفتوح

OPTS-TTPO: Enhancing Finite-Sample Policy-Gradient Learning with Tree Search

Junyu Lu, Shichao Weng, Zhiqiang Wang وآخرون · 2026

The policy-gradient theorem gives the exact gradient under the current policy, but finite on-policy samples may miss rare high-return trajectories. We study whether tree search improves their coverage within a fixed budget while controlling gradient bias. We introduce On-Policy Parallel Tree Search (OPTS) and Tree Traj …

المؤلفون المشاركون