[
    {
        "id": "osp-24036",
        "type": "article-journal",
        "title": "GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking",
        "author": [
            {
                "family": "Liu",
                "given": "Jian"
            },
            {
                "family": "Sun",
                "given": "Wei"
            },
            {
                "family": "Dai",
                "given": "Zhenqi"
            },
            {
                "family": "Yang",
                "given": "Hui"
            },
            {
                "family": "Xiao",
                "given": "Jian"
            },
            {
                "family": "Sebe",
                "given": "Nicu"
            },
            {
                "family": "Zhao",
                "given": "Na"
            }
        ],
        "URL": "https://omanscience.com/en/articles/gencope-syn2real-generalized-category-level-object-pose-estimation-for-robotic-picking",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Category-level object pose estimation (COPE), capable of generalizing to intra-class unknown objects, has become a core technique for robotic 3D scene understanding. However, existing COPE methods still require labor-intensive recollection of real-world training data for novel object categories, which limits their scalability in practical applications. This paper aims to achieve synthetic-to-real (Syn2Real) generalized COPE, where a model is trained solely on rendered synthetic data and directly generalized to real-world deployments. The central challenge lies in the significant domain gap between synthetic and real-world data, particularly in texture appearance. To address this, we aim to enhance domain generalization by learning domain-invariant representations that capture semantic commonalities among objects within the same category. We introduce 2D and 3D semantic consistency constraints to reduce the sensitivity of feature encoders to domain-specific features. In addition, we propose an end-to-end pose regression framework that performs 2D-3D cross consistency learning, leveraging dense cross-modality fusion to further refine pose estimation. Since simplicity and effectiveness are essential for real-world robotic deployment, our model operates exclusively on global features, yielding a highly lightweight and efficient architecture. Extensive experiments on the REAL275 and Wild6D benchmarks, as well as real-world robotic manipulation scenes, show superior Syn2Real generalization performance of our paradigm. Code and demos are released at https://paperreview99.github.io/GenCOPE/."
    }
]