[
    {
        "id": "osp-25634",
        "type": "article-journal",
        "title": "T$^2$Mem: Learning Test-Time Memory for Robotics",
        "author": [
            {
                "family": "Liu",
                "given": "Yize"
            },
            {
                "family": "Huang",
                "given": "Huang"
            },
            {
                "family": "Hong",
                "given": "Yining"
            },
            {
                "family": "Du",
                "given": "Zijian"
            },
            {
                "family": "Cao",
                "given": "Zhi"
            },
            {
                "family": "Fei-Fei",
                "given": "Li"
            },
            {
                "family": "Wu",
                "given": "Jiajun"
            }
        ],
        "URL": "https://omanscience.com/ar/articles/t-2-mem-learning-test-time-memory-for-robotics",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Memory-dependent robotic manipulation requires policies to use information that is no longer available in the current observation. Retaining history alone is insufficient: memory must preserve information that supports future actions. One challenge is whether a memory-free foundation model can learn to retain and use historical information from action demonstrations alone, without external memory support. We introduce T$^2$Mem, a framework that develops this capability within a pretrained vision-language-action policy, without external reasoning models or memory-specific annotations. T$^2$Mem uses test-time training to encode observation history into compact fast weights through online self-supervised updates, avoiding repeated processing of the full history. An observation-grounded interface extracts vision-language information for memory formation and supplies retrieved context to the action expert. Action supervision shapes what the memory learns to retain and use, while alternating memory-policy learning gives each component a fixed counterpart during optimization. Across 16 RoboMME tasks, T$^2$Mem improves average success from 17.93% to 56.83% over the memory-free base policy and outperforms the recurrent-memory methods reported in the benchmark, while controlled profiling indicates at least 3x inference speedup over explicit methods. Project website: https://yzliu84.github.io/T2MEM-project/"
    }
]