الملخص
Video generation must account for two sources of motion, one induced by the observer's camera path and the other caused by scene dynamics. An ideal camera-controlled video model should account for both motions: let users move the camera while evolving the scene dynamics. While current models handle camera-induced motion well in static settings, they struggle for dynamic scenes: objects are static, move incorrectly, or degrade in generation quality. We introduce DynaTokens, a lightweight set of learnable scene-specific tokens that teach dynamics to an existing camera-controlled world model. Our method is motivated by a simple asymmetry between the two sources of motion: whereas camera motion affects the generated view globally, object dynamics are spatially localized. Through cross-attention, DynaTokens trains the learnable tokens from a few example trajectories for a scene while keeping the base model frozen, and enables dynamics under new query camera paths. DynaTokens achieves a better simultaneous dynamics-camera tradeoff on VBench2 and WorldScore evaluations than LoRA, block finetuning, and specialized trainable-layer baselines. Analyses of token attention, ablations, and motion temporality suggest that matching the trainable interface to the structure of the learning target is important for effective adaptation. Project website: https://glab-caltech.github.io/dynatokens/
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Ma, Z., Chen, H., & Gkioxari, G. (2026). DynaTokens: Teaching Dynamics to Camera-Controlled Video Models at Test Time. https://omanscience.com/ar/articles/dynatokens-teaching-dynamics-to-camera-controlled-video-models-at-test-time
MLA 9
Ma, Ziqi, et al. "DynaTokens: Teaching Dynamics to Camera-Controlled Video Models at Test Time." https://omanscience.com/ar/articles/dynatokens-teaching-dynamics-to-camera-controlled-video-models-at-test-time.
شيكاغو (المؤلف–التاريخ)
Ma, Ziqi, Hongqiao Chen, and Georgia Gkioxari. 2026. "DynaTokens: Teaching Dynamics to Camera-Controlled Video Models at Test Time." https://omanscience.com/ar/articles/dynatokens-teaching-dynamics-to-camera-controlled-video-models-at-test-time.
هارفارد
Ma, Z., Chen, H. and Gkioxari, G. (2026) 'DynaTokens: Teaching Dynamics to Camera-Controlled Video Models at Test Time', Available at: https://omanscience.com/ar/articles/dynatokens-teaching-dynamics-to-camera-controlled-video-models-at-test-time.
فانكوفر
Ma Z, Chen H, Gkioxari G. DynaTokens: Teaching Dynamics to Camera-Controlled Video Models at Test Time. https://omanscience.com/ar/articles/dynatokens-teaching-dynamics-to-camera-controlled-video-models-at-test-time
IEEE
Z. Ma, H. Chen, and G. Gkioxari, "DynaTokens: Teaching Dynamics to Camera-Controlled Video Models at Test Time," https://omanscience.com/ar/articles/dynatokens-teaching-dynamics-to-camera-controlled-video-models-at-test-time.