الملخص

Generating high-quality and controllable camera motion is essential for AI-assisted cinematography, video synthesis, and 3D scene understanding. We introduce TKCAM, a Text- and Keyframe-conditioned CAMera-motion synthesis framework based on generative masked modeling. We represent camera dynamics using a 12-dimensional kinematic feature comprising position, velocity, and a continuous rotation representation and discretize them into hierarchical motion tokens via a Residual Vector Quantizer (RVQ). A two-stage masked transformer architecture then learns to reconstruct and refine these tokens, utilizing explicit self- and cross-attention modules for multimodal conditioning. A central feature of our framework is sparse visual keyframe conditioning: users can provide free-form text prompts together with RGB observations at selected timestamps, which provide temporally localized visual guidance for generating coherent in-between trajectories. Furthermore, to advance evaluation standards, we curate RealEstate10K-Cap, a large-scale text-camera dataset, and establish a cross-domain benchmark with a Universal CLaTr Evaluator. Extensive experiments demonstrate that TKCAM surpasses recent state-of-the-art baselines on Fréchet distance (FID), text-motion matching scores, and retrieval metrics (R@K), while additional analyses evaluate temporal smoothness and cross-domain generalization. Code is available at https://github.com/linearalgebrayhz/TKCAM.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Yang, H., Dou, Z., Gu, Z., Lin, C., Wang, W., Liu, Y., & Komura, T. (2026). TKCAM: Text and Keyframe to Camera Trajectory Generation. https://omanscience.com/ar/articles/tkcam-text-and-keyframe-to-camera-trajectory-generation

MLA 9

Yang, Haozhe, et al. "TKCAM: Text and Keyframe to Camera Trajectory Generation." https://omanscience.com/ar/articles/tkcam-text-and-keyframe-to-camera-trajectory-generation.

شيكاغو (المؤلف–التاريخ)

Yang, Haozhe, Zhiyang Dou, Zekai Gu, Cheng Lin, Wenping Wang, Yuan Liu, and Taku Komura. 2026. "TKCAM: Text and Keyframe to Camera Trajectory Generation." https://omanscience.com/ar/articles/tkcam-text-and-keyframe-to-camera-trajectory-generation.

هارفارد

Yang, H., Dou, Z., Gu, Z., Lin, C., Wang, W., Liu, Y. and Komura, T. (2026) 'TKCAM: Text and Keyframe to Camera Trajectory Generation', Available at: https://omanscience.com/ar/articles/tkcam-text-and-keyframe-to-camera-trajectory-generation.

فانكوفر

Yang H, Dou Z, Gu Z, Lin C, Wang W, Liu Y, et al. TKCAM: Text and Keyframe to Camera Trajectory Generation. https://omanscience.com/ar/articles/tkcam-text-and-keyframe-to-camera-trajectory-generation

IEEE

H. Yang, Z. Dou, Z. Gu, C. Lin, W. Wang, Y. Liu, and T. Komura, "TKCAM: Text and Keyframe to Camera Trajectory Generation," https://omanscience.com/ar/articles/tkcam-text-and-keyframe-to-camera-trajectory-generation.