Abstract

Training embodied foundation models typically requires massive-scale datasets and extensive computational resources, yet often suffers from three critical limitations: (1) inefficient sample utilization due to low-informative samples; (2) imbalanced gradient contributions across heterogeneous tasks; and (3) severe credit assignment problem in long-horizon planning, where trajectory-level rewards indiscriminately penalize all tokens. To address these issues, we propose an efficient training paradigm that achieves state-of-the-art average performance through strategic data selection and hierarchical policy optimization. Our approach consists of three synergistic stages. First, Rejection Sampling-based Fine-Tuning (RSFT) filters out low-informative samples to establish robust behavioral priors while preventing distributional collapse. Second, Iterative Rejection GRPO (IR-GRPO) employs task-specific queues stratified by difficulty to keep datasets balanced across reinforcement learning iterations, coupled with a hybrid reward mechanism for precise cross-task feedback. Third, to enhance long-horizon task planning, we introduce Trie-GRPO, a novel reinforcement learning algorithm based on action prefix trees, which enables step-level advantage estimation. This resolves the credit assignment problem by isolating intermediate correct decisions from downstream errors, while effectively balancing exploration efficiency and depth compared to conventional search trees. As a result, EmbodiedMind achieves a state-of-the-art average performance of 70.02% across 18 benchmarks, and significantly outperforms other embodied foundation models in long-horizon task planning accuracy. Our project will be released for reproducibility.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Wang, F., Zhang, Z., Zhang, Y., Wang, L., Zhu, Y., Deng, J., Zhang, M., Gao, Z., Wang, Y., Xu, J., & Yang, R. (2026). EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence. https://omanscience.com/en/articles/embodiedmind-adaptive-data-curation-and-prefix-tree-reinforcement-learning-for-efficient-embodied-intelligence

MLA 9

Wang, Feifan, et al. "EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence." https://omanscience.com/en/articles/embodiedmind-adaptive-data-curation-and-prefix-tree-reinforcement-learning-for-efficient-embodied-intelligence.

Chicago (author–date)

Wang, Feifan, Zongbing Zhang, Yu Zhang, Lingfeng Wang, Yurui Zhu, Jin Deng, Mingliang Zhang, Zhengguang Gao, Yongcheng Wang, Jin Xu, and Ri Yang. 2026. "EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence." https://omanscience.com/en/articles/embodiedmind-adaptive-data-curation-and-prefix-tree-reinforcement-learning-for-efficient-embodied-intelligence.

Harvard

Wang, F., Zhang, Z., Zhang, Y., Wang, L., Zhu, Y., Deng, J., Zhang, M., Gao, Z., Wang, Y., Xu, J. and Yang, R. (2026) 'EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence', Available at: https://omanscience.com/en/articles/embodiedmind-adaptive-data-curation-and-prefix-tree-reinforcement-learning-for-efficient-embodied-intelligence.

Vancouver

Wang F, Zhang Z, Zhang Y, Wang L, Zhu Y, Deng J, et al. EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence. https://omanscience.com/en/articles/embodiedmind-adaptive-data-curation-and-prefix-tree-reinforcement-learning-for-efficient-embodied-intelligence

IEEE

F. Wang, Z. Zhang, Y. Zhang, L. Wang, Y. Zhu, J. Deng, M. Zhang, Z. Gao, Y. Wang, J. Xu, and R. Yang, "EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence," https://omanscience.com/en/articles/embodiedmind-adaptive-data-curation-and-prefix-tree-reinforcement-learning-for-efficient-embodied-intelligence.