[
    {
        "id": "osp-24130",
        "type": "article-journal",
        "title": "StreamRig: Exploiting Intra-Rig Geometry for Streaming Multi-Camera Odometry",
        "author": [
            {
                "family": "Wei",
                "given": "Yufei"
            },
            {
                "family": "Ye",
                "given": "Shuhao"
            },
            {
                "family": "Wang",
                "given": "Qi"
            },
            {
                "family": "Zheng",
                "given": "Xin"
            },
            {
                "family": "Huang",
                "given": "Qing"
            },
            {
                "family": "Xiong",
                "given": "Rong"
            },
            {
                "family": "Wang",
                "given": "Yue"
            }
        ],
        "URL": "https://omanscience.com/en/articles/streamrig-exploiting-intra-rig-geometry-for-streaming-multi-camera-odometry",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving efficient use of rig geometry a challenge. We present StreamRig, a freeze-and-stream framework that builds causal streaming odometry for calibrated rigs on a frozen multi-view 3D foundation model. The frozen front-end jointly perceives the synchronized views using rig calibration. A Rig-Resampler compresses their features, a CausalBridge applies causal attention with a key-value cache, and a lightweight head regresses rig poses. A periodic re-anchoring protocol supports stable pose estimation over long sequences. Only these modules are trained, 74.6M parameters in total, with relative poses as the sole supervision. Our two-stage training strategy combines group relocalization pretraining with causal rig training to transfer the geometric priors of the frozen front-end and the alignment ability of the pretrained modules to streaming odometry. We evaluate on NCLT, TartanGround, KITTI-360, and our self-collected humanoid-robot dataset ZJH, where training uses only simulation and real-world evaluation is zero-shot. Across all four datasets, StreamRig achieves lower translation and rotation drift than the evaluated non-oracle monocular streaming and rig-aware offline models, while maintaining low inference cost. Ablations and controlled camera-count experiments identify the sources of these gains. We further examine how longer training windows affect inference over longer horizons. Code has been released at https://github.com/WeiYuFei0217/StreamRig."
    }
]