الباحثون

Xiaomin Li

المنشورات 4

نسخة أولية وصول مفتوح

Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts

Multi-teacher on-policy distillation (MOPD) is used in two settings. In common-domain composition, several teachers score each student rollout from one prompt domain and their signals form a single target; in routed-domain distillation, prompts from different domains are assigned to the corresponding specialist. Both s …

نسخة أولية وصول مفتوح

AgentPersonaBench: Benchmarking Persona-Driven User Simulation

Jintao Huang, Yifan Wang, Hongyu Shen وآخرون · 2026

We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic b …

نسخة أولية وصول مفتوح

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement

Fang Wu, Da Xing, Yanjie Huang وآخرون · 2026

Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback …

نسخة أولية وصول مفتوح

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed …

المؤلفون المشاركون