[
    {
        "id": "osp-25851",
        "type": "article-journal",
        "title": "DualManip: Agentic Dynamic Manipulation via Dual-Path Semantic Reasoning and Geometric Adaptation",
        "author": [
            {
                "family": "Li",
                "given": "Chengxi"
            },
            {
                "family": "Di",
                "given": "Yan"
            },
            {
                "family": "Li",
                "given": "Yingyue"
            },
            {
                "family": "Zhang",
                "given": "Ruida"
            },
            {
                "family": "Li",
                "given": "Mingyang"
            },
            {
                "family": "Ji",
                "given": "Xiangyang"
            }
        ],
        "URL": "https://omanscience.com/en/articles/dualmanip-agentic-dynamic-manipulation-via-dual-path-semantic-reasoning-and-geometric-adaptation",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Vision-language models (VLMs) enable open-vocabulary reasoning for robot manipulation, but their high inference latency limits responsiveness in dynamic scenes. Many scene changes, however, alter object geometry without invalidating task intent. We present DualManip, a dual-path framework that decouples infrequent semantic reasoning from responsive geometric adaptation. The semantic path decomposes the task and grounds task-relevant interactions, followed by a constraint-solving module for pose optimization. During execution, the geometric path continuously updates template-to-observation correspondences from live RGB-D observations via a shape-adaptive network. These correspondences transfer task-relevant grasp contacts across observations, enabling online grasp reconstruction under object motion and non-rigid deformation. The Information Interaction Module bridges the two paths by initializing task-relevant grasps from semantic grounding, validating geometric updates, and triggering semantic replanning upon update failures. Real-world evaluation spans six manipulation tasks covering non-rigid deformation, articulated reconfiguration, rigid motion, and high-precision assembly across three settings: static, single-change, and continuous dynamic. DualManip demonstrates superior manipulation robustness, particularly under continuous scene changes, while achieving geometric adaptation approximately 46$\\times$ faster than agentic verification and semantic replanning. Our project page: https://lichengxi1.github.io/Dualmanip."
    }
]