الباحثون

Wenbin Hu

المنشورات 2

نسخة أولية وصول مفتوح

SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown

Yuxuan Liu, Haoran Li, Yuhao Zhang وآخرون · 2026

Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish documentation repair, script repair, and preservation when evaluating skill s …

نسخة أولية وصول مفتوح

CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning

Wenbin Hu, Huihao Jing, Haochen Shi وآخرون · 2026

Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation. …

المؤلفون المشاركون