Authors

Daoan Zhang

Publications 5

Preprint Open access

LIBERO-MAX: Do Robot Policies Adapt When the World Changes?

Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introdu …

Preprint Open access

Beyond Oracle Communication: Benchmarking Interactive Intent Alignment Under Miscommunication and Evolving User Intent

Modern LLM agents increasingly tackle complex tasks through interactive, long-horizon exchanges with users, while existing benchmarks generally assume that users always accurately and sufficiently communicate a fixed intent. However, this oracle communication assumption rarely holds in practice: users may miscommunicat …

Preprint Open access

TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

Dehai Min, Daoan Zhang, Yiming Zeng et al. · 2026

An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable …

Co-authors