[
    {
        "id": "osp-24131",
        "type": "article-journal",
        "title": "LOCI: Spatial Linear Memory for Streaming World Models",
        "author": [
            {
                "family": "Xia",
                "given": "Ji"
            },
            {
                "family": "Liao",
                "given": "Tingting"
            },
            {
                "family": "Liang",
                "given": "Xuezhi"
            },
            {
                "family": "Li",
                "given": "Hao"
            },
            {
                "family": "Liu",
                "given": "Guangyi"
            }
        ],
        "URL": "https://omanscience.com/en/articles/loci-spatial-linear-memory-for-streaming-world-models",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but compresses history into a fixed-size state, so individual past observations are no longer directly accessible. We introduce LOCI, a hybrid spatial-memory architecture that keeps both representations. In half of the transformer blocks, main attention keeps a key-value cache of past observations; in the other half, it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, so viewpoint enters both memory addressing and stored content. Recurrent readouts flow into subsequent cache-backed blocks and supply their queries with accumulated scene context. On the public MIND memory benchmark and on held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model; with full history, it lowers peak memory at equal length by about 30% relative to full softmax. With a bounded bank of retained observations, it streams long videos at constant memory and remains more faithful than full softmax under the same budget."
    }
]