Abstract
Camera trajectories control viewpoint changes in video generation, scene reconstruction, and robotic perception. Generating them from language requires both scene geometry and target-aware framing. We introduce OmniCam, an autoregressive model that generates camera pose sequences from a single panorama and textual trajectory descriptions. Its geometry-grounded pose token learning combines three components: a panoramic point-cloud encoder for omnidirectional geometric context; hybrid absolute-rotation and relative-translation tokenization with temporally consistent quaternion signs; and separate geometric and semantic conditioning streams with an explicit 3D target anchor. We also construct OmniCaT, containing 267,700 trajectories across four camera behaviors. On the reported OmniCaT evaluation, OmniCam reduces trajectory errors by 28--47% and collision rate by 65.8% relative to GenDoP retrained on OmniCaT. Against the best baseline for each metric, the ATE and collision reductions are 43.0% and 62.3%, respectively. Component ablations support the use of geometric and target-aware conditioning, while downstream experiments examine camera-controlled video generation and robotic active perception.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Liu, Z., Cao, C., Zhang, Y., Zuo, X., Xue, X., Fu, Y., Wang, T., & Guo, C. (2026). OmniCam: Omni-Camera Trajectory Generation via Geometry-Grounded Pose Token Learning. https://omanscience.com/en/articles/omnicam-omni-camera-trajectory-generation-via-geometry-grounded-pose-token-learning
MLA 9
Liu, Zhenyang, et al. "OmniCam: Omni-Camera Trajectory Generation via Geometry-Grounded Pose Token Learning." https://omanscience.com/en/articles/omnicam-omni-camera-trajectory-generation-via-geometry-grounded-pose-token-learning.
Chicago (author–date)
Liu, Zhenyang, Chenjie Cao, Yisu Zhang, Xuhui Zuo, Xiangyang Xue, Yanwei Fu, Tengfei Wang, and Chunchao Guo. 2026. "OmniCam: Omni-Camera Trajectory Generation via Geometry-Grounded Pose Token Learning." https://omanscience.com/en/articles/omnicam-omni-camera-trajectory-generation-via-geometry-grounded-pose-token-learning.
Harvard
Liu, Z., Cao, C., Zhang, Y., Zuo, X., Xue, X., Fu, Y., Wang, T. and Guo, C. (2026) 'OmniCam: Omni-Camera Trajectory Generation via Geometry-Grounded Pose Token Learning', Available at: https://omanscience.com/en/articles/omnicam-omni-camera-trajectory-generation-via-geometry-grounded-pose-token-learning.
Vancouver
Liu Z, Cao C, Zhang Y, Zuo X, Xue X, Fu Y, et al. OmniCam: Omni-Camera Trajectory Generation via Geometry-Grounded Pose Token Learning. https://omanscience.com/en/articles/omnicam-omni-camera-trajectory-generation-via-geometry-grounded-pose-token-learning
IEEE
Z. Liu, C. Cao, Y. Zhang, X. Zuo, X. Xue, Y. Fu, T. Wang, and C. Guo, "OmniCam: Omni-Camera Trajectory Generation via Geometry-Grounded Pose Token Learning," https://omanscience.com/en/articles/omnicam-omni-camera-trajectory-generation-via-geometry-grounded-pose-token-learning.