Abstract

Achieving robust cross-embodiment generalization in imitation learning demands overcoming a critical representation flaw that inextricably entangles task semantics with hardware-specific visual geometry. We propose an interaction-centric framework that leverages the shared structure of two-finger grippers via a parameterized universal gripper abstraction, yielding a canonical gripper-frame representation. Given language and RGB-D observations, a VLM infers the subtask and grounds an interaction triplet (gripper, held, target), while SAM~2.1 tracks masks to reduce VLM queries. We design concise hybrid features that combine target/collision artificial potential fields for global guidance with segmented gripper-frame point clouds for local geometry, and use a Flow-Matching Transformer to predict smooth 7-DoF action chunks. Experiments in simulation and real-world tasks demonstrate that ours is the first imitation learning approach to simultaneously achieve competitive benchmark scores and extreme cross-embodiment/cross-viewpoint zero-shot sim-to-real transfer to completely distinct, heterogeneous robot platforms.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Li, G., Bao, S., Zhao, Y., Shen, H., Li, H., Zhao, C., Yang, T., Tang, J., & Zhang, J. (2026). Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling. https://omanscience.com/en/articles/enabling-a-unified-cross-domain-representation-for-two-finger-gripper-manipulation-via-interaction-centric-modeling

MLA 9

Li, Guanlin, et al. "Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling." https://omanscience.com/en/articles/enabling-a-unified-cross-domain-representation-for-two-finger-gripper-manipulation-via-interaction-centric-modeling.

Chicago (author–date)

Li, Guanlin, Shifeng Bao, Yihan Zhao, Haitao Shen, Haoyang Li, Chen Zhao, Tong Yang, Jie Tang, and Jing Zhang. 2026. "Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling." https://omanscience.com/en/articles/enabling-a-unified-cross-domain-representation-for-two-finger-gripper-manipulation-via-interaction-centric-modeling.

Harvard

Li, G., Bao, S., Zhao, Y., Shen, H., Li, H., Zhao, C., Yang, T., Tang, J. and Zhang, J. (2026) 'Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling', Available at: https://omanscience.com/en/articles/enabling-a-unified-cross-domain-representation-for-two-finger-gripper-manipulation-via-interaction-centric-modeling.

Vancouver

Li G, Bao S, Zhao Y, Shen H, Li H, Zhao C, et al. Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling. https://omanscience.com/en/articles/enabling-a-unified-cross-domain-representation-for-two-finger-gripper-manipulation-via-interaction-centric-modeling

IEEE

G. Li, S. Bao, Y. Zhao, H. Shen, H. Li, C. Zhao, T. Yang, J. Tang, and J. Zhang, "Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling," https://omanscience.com/en/articles/enabling-a-unified-cross-domain-representation-for-two-finger-gripper-manipulation-via-interaction-centric-modeling.