[
    {
        "id": "osp-15956",
        "type": "article-journal",
        "title": "OGAM: Connecting Systematic Testing to Runtime Assurance through Object-Grounded Attention Monitoring for VLA Policies",
        "author": [
            {
                "family": "Darwish",
                "given": "Haki"
            },
            {
                "family": "Yin",
                "given": "Xiangyu"
            },
            {
                "family": "Li",
                "given": "Changwen"
            },
            {
                "family": "Yan",
                "given": "Rongjie"
            },
            {
                "family": "Neto",
                "given": "Francisco Gomes de Oliveira"
            },
            {
                "family": "Cheng",
                "given": "Chih-Hong"
            }
        ],
        "URL": "https://omanscience.com/en/articles/ogam-connecting-systematic-testing-to-runtime-assurance-through-object-grounded-attention-monitoring-for-vla-policies",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Benchmarks expose vision-language-action (VLA) policies to few canonical instructions, while exhaustive deployment testing is impossible. We introduce Object-Grounded Attention Monitoring (OGAM), connecting systematic testing to runtime assurance: testing reveals attention divergence between successful and failed executions, and OGAM uses this signal to stop failures beyond the finite suite. We generate scene-grounded instructions through pairwise combinations of action templates and objects, and separately test meaning-preserving paraphrases. All 87 out-of-benchmark cases reveal problematic behavior across OpenVLA, OpenVLA-OFT, UniVLA, and $π_{0.5}$: none completes any of the 24 feasible instructions, while infeasible or hazardous requests also trigger behavior substitution. At each action query, we project gradient-weighted visual attention through object masks and group it by instruction role for comparison across tasks and policies. Dynamic time warping aligns this course with a successful reference despite speed differences; conformal calibration on successful episodes sets the early-stopping threshold for sustained deviations, with a nominal false-stop target of $α=0.05$. Across four policies, OGAM stops 87-100% of failed episodes at median times of 5-12s within a 20s budget, with observed false-stop rates of 3-5%, without failure-labeled training. Finite testing thus identifies attention patterns that support online intervention before failure fully unfolds."
    }
]