الباحثون

Daoan Zhang

المنشورات 5

نسخة أولية وصول مفتوح

Dependency-Aware Reward Shaping for Agentic Reinforcement Learning

Ziyi Chen, Yan Zhang, Jianhui Wei وآخرون · 2026

When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted while independent work remains valid. With only a final success/failure reward, every st …

نسخة أولية وصول مفتوح

Complementary Retrieval-Augmented Prompting for Consistent Long-Form Video Generation

Xianghan Wei, Xiaoda Yang, Zhi Wang وآخرون · 2026

While recent video foundation models excel at generating high-quality short videos, long-form video generation remains a critical challenge, where a major bottleneck lies in conditioning independently generated shots to preserve consistent characters, scenes, and objects throughout a story. Existing training-free appro …

نسخة أولية وصول مفتوح

LIBERO-MAX: Do Robot Policies Adapt When the World Changes?

Yunbei Zhang, Zijian Jin, Yuanzhe Liu وآخرون · 2026

Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introdu …

نسخة أولية وصول مفتوح

Beyond Oracle Communication: Benchmarking Interactive Intent Alignment Under Miscommunication and Evolving User Intent

Zheyuan Zhang, Mengyuan Chao, Ke Xiao وآخرون · 2026

Modern LLM agents increasingly tackle complex tasks through interactive, long-horizon exchanges with users, while existing benchmarks generally assume that users always accurately and sufficiently communicate a fixed intent. However, this oracle communication assumption rarely holds in practice: users may miscommunicat …

نسخة أولية وصول مفتوح

TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

Dehai Min, Daoan Zhang, Yiming Zeng وآخرون · 2026

An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable …

المؤلفون المشاركون