[
    {
        "id": "osp-24929",
        "type": "article-journal",
        "title": "Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling",
        "author": [
            {
                "family": "Li",
                "given": "Guanlin"
            },
            {
                "family": "Bao",
                "given": "Shifeng"
            },
            {
                "family": "Zhao",
                "given": "Yihan"
            },
            {
                "family": "Shen",
                "given": "Haitao"
            },
            {
                "family": "Li",
                "given": "Haoyang"
            },
            {
                "family": "Zhao",
                "given": "Chen"
            },
            {
                "family": "Yang",
                "given": "Tong"
            },
            {
                "family": "Tang",
                "given": "Jie"
            },
            {
                "family": "Zhang",
                "given": "Jing"
            }
        ],
        "URL": "https://omanscience.com/ar/articles/enabling-a-unified-cross-domain-representation-for-two-finger-gripper-manipulation-via-interaction-centric-modeling",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Achieving robust cross-embodiment generalization in imitation learning demands overcoming a critical representation flaw that inextricably entangles task semantics with hardware-specific visual geometry. We propose an interaction-centric framework that leverages the shared structure of two-finger grippers via a parameterized universal gripper abstraction, yielding a canonical gripper-frame representation. Given language and RGB-D observations, a VLM infers the subtask and grounds an interaction triplet (gripper, held, target), while SAM~2.1 tracks masks to reduce VLM queries. We design concise hybrid features that combine target/collision artificial potential fields for global guidance with segmented gripper-frame point clouds for local geometry, and use a Flow-Matching Transformer to predict smooth 7-DoF action chunks. Experiments in simulation and real-world tasks demonstrate that ours is the first imitation learning approach to simultaneously achieve competitive benchmark scores and extreme cross-embodiment/cross-viewpoint zero-shot sim-to-real transfer to completely distinct, heterogeneous robot platforms."
    }
]