الباحثون

Lan-Zhe Guo

المنشورات 4

نسخة أولية وصول مفتوح

T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning

Bowen Zhang, Junwei He, Maoqi Liu وآخرون · 2026

Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajecto …

نسخة أولية وصول مفتوح

HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness

Yang Chen, Lirong Che, Zhenyu Huang وآخرون · 2026

Embodied navigation requires agents to ground instructions or object goals in spatial observations and translate plans into successful execution. As multimodal large language models (MLLMs) become increasingly capable, they offer stronger support for navigation without task-specific training; however, improved semantic …

المؤلفون المشاركون