الباحثون

Song-Lin Lv

المنشورات 1

نسخة أولية وصول مفتوح

T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning

Bowen Zhang, Junwei He, Maoqi Liu وآخرون · 2026

Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajecto …

المؤلفون المشاركون