الباحثون

Pengxin Wang

المنشورات 2

نسخة أولية وصول مفتوح

Learning Perturbation Robust Policies for LLM Agents with Stable Optimization

Pengxin Wang, Yuanzhe Li, Yuxin Ren وآخرون · 2026

Reinforcement learning (RL) has become an effective post-training paradigm for long-horizon large language model (LLM) agents. However, we find that the resulting policies can be sensitive to various policy perturbations, such as hidden-state noise, pruning, and quantization. In this work, we study how to improve pertu …

المؤلفون المشاركون