Abstract

Long-horizon mobile manipulation requires a robot to navigate multi-room environments and execute a sequence of manipulation skills under a single natural language instruction. Learning and evaluating these skills present three challenges: similar observations under a fixed task instruction may make skill selection ambiguous; even when a preceding skill succeeds, the robot state inherited by the next skill may deviate from its demonstrated starting states and affect execution; and task-level metrics hinder skill-specific diagnosis, while early failures leave later skills untested. We therefore introduce ManiUnit, a manipulation skill dataset and benchmark built from 50 BEHAVIOR-1K activities. Its dataset contains 137,899 segments across 21 skill types and 417 subtasks, and its benchmark contains 1,260 test instances. Correspondingly, ManiUnit pairs each segment with an explicit subtask instruction; measures sensitivity to perturbations of the robot's starting base position or joint configuration; and restores intermediate simulator states and defines local success conditions so that each skill can be evaluated without executing preceding stages. Evaluations of representative vision-language-action (VLA) policies show that similar aggregate scores can hide substantial per-skill differences. The tested starting-state perturbations also degrade execution: on the full benchmark, joint perturbations reduce success rates by approximately 56% relative to those from demonstrated starting states. On two long-horizon activities, a skill policy trained on ManiUnit segments achieves 78.7% local manipulation success, compared with 49.3% for a task policy trained on complete demonstrations. The trained skills further support complete-task execution on these activities, as coordinating the task and skill policies through a planner raises full-task success from 4.0% to 18.0%.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Wei, G., Yan, D., Yuan, X., Zhang, G., He, Y., Du, G., Ye, J., Zhang, H., Wei, X., Qi, X., Zhao, C., Zhang, H., & Xiao, R. (2026). ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks. https://omanscience.com/en/articles/maniunit-a-manipulation-skill-dataset-and-benchmark-for-long-horizon-tasks

MLA 9

Wei, Guoting, et al. "ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks." https://omanscience.com/en/articles/maniunit-a-manipulation-skill-dataset-and-benchmark-for-long-horizon-tasks.

Chicago (author–date)

Wei, Guoting, Dawei Yan, Xia Yuan, Gengming Zhang, Yelin He, Guodong Du, Jiaquan Ye, Heng Zhang, Xinming Wei, Xianbiao Qi, Chunxia Zhao, Haokui Zhang, and Rong Xiao. 2026. "ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks." https://omanscience.com/en/articles/maniunit-a-manipulation-skill-dataset-and-benchmark-for-long-horizon-tasks.

Harvard

Wei, G., Yan, D., Yuan, X., Zhang, G., He, Y., Du, G., Ye, J., Zhang, H., Wei, X., Qi, X., Zhao, C., Zhang, H. and Xiao, R. (2026) 'ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks', Available at: https://omanscience.com/en/articles/maniunit-a-manipulation-skill-dataset-and-benchmark-for-long-horizon-tasks.

Vancouver

Wei G, Yan D, Yuan X, Zhang G, He Y, Du G, et al. ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks. https://omanscience.com/en/articles/maniunit-a-manipulation-skill-dataset-and-benchmark-for-long-horizon-tasks

IEEE

G. Wei, D. Yan, X. Yuan, G. Zhang, Y. He, G. Du, J. Ye, H. Zhang, X. Wei, X. Qi, C. Zhao, H. Zhang, and R. Xiao, "ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks," https://omanscience.com/en/articles/maniunit-a-manipulation-skill-dataset-and-benchmark-for-long-horizon-tasks.