الباحثون

Yuexing Hao

المنشورات 3

نسخة أولية وصول مفتوح

AgentPersonaBench: Benchmarking Persona-Driven User Simulation

Jintao Huang, Yifan Wang, Hongyu Shen وآخرون · 2026

We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic b …

نسخة أولية وصول مفتوح

Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts

As large language models (LLMs) are increasingly used in clinical settings, it is critical to evaluate their reliability under realistic variation in clinical text. We study this question in clinical triage, comparing LLMs to practicing physicians under text perturbations that preserve the underlying clinical setting. …

نسخة أولية وصول مفتوح

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed …

المؤلفون المشاركون