[
    {
        "id": "osp-26285",
        "type": "article-journal",
        "title": "Selective Commitment for Language-Guided Object Retrieval under Partial Observability",
        "author": [
            {
                "family": "Koh",
                "given": "Wonhee"
            },
            {
                "family": "Dinesh",
                "given": "Sushil Samuel"
            },
            {
                "family": "Ko",
                "given": "Hansol"
            },
            {
                "family": "Park",
                "given": "Shinkyu"
            },
            {
                "family": "Lee",
                "given": "Eungjoo"
            }
        ],
        "URL": "https://omanscience.com/en/articles/selective-commitment-for-language-guided-object-retrieval-under-partial-observability",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Language-guided object retrieval under partial observability requires deciding whether to gather more evidence, interact with the scene, grasp a candidate, or abstain. We present a closed-loop framework that coordinates these decisions for retrieving a target specified in relation to a reference container. The framework maintains a persistent joint belief over target identity, container relation, and presence through tracked-object, unobserved-target, and target-absent hypotheses. View-conditioned categorical VLM observations update this belief; conformal grasp eligibility and robot feasibility govern commitment, while finite-horizon belief-space planning selects information-gathering actions. Across five different scenarios, our proposed method succeeds in 19/25 simulation episodes versus 12/25 for the best-performing task-adapted baseline and is the only evaluated policy to achieve at least one success in each scenario. Ablations show that cross-view memory improves success under partial occlusion, while the full system does not consistently outperform simplified variants. Real-robot trials demonstrate closed-loop re-observation and autonomous recovery from injected grasp failures, while injected viewpoint failures end in false defer. Experimental results demonstrate the feasibility of coordinating evidence gathering and selective grasp commitment within a unified framework for retrieval under partial observability."
    }
]