الملخص
Large generative models can recover realistic detail in real-world video super-resolution (VSR), but processing an entire video with them is computationally expensive. In this work, we present RelayVSR, a streaming VSR framework built on the Sparse Generative Relay mechanism. A large generative model generates reference latents for sparse keyframes, while a lightweight VSR network uses these references and low-resolution video to super-resolve every frame. The lightweight VSR network, implemented as a Dual-Memory Video Transformer, reuses keyframe information across frames and updates recent video context, supporting first-keyframe conditioning and dual-endpoint conditioning with bounded lookahead. However, errors in shared keyframes can propagate and accumulate across output frames, making keyframe quality alone an insufficient optimization target. We address this collaboration gap with Video-Aware Reference Optimization (VARO), which uses reinforcement learning to update the large generative model with two reward levels: a system-level reward evaluates videos produced by the fixed lightweight VSR network, while a reference-level reward evaluates decoded keyframe quality. VARO improves final video quality over direct joint training, and its dual-level rewards outperform a system-level reward alone. At 1080p on a single NVIDIA A100 80GB, dual-endpoint RelayVSR with a 15-frame keyframe interval reaches 29.29 FPS, 13.82 GB peak GPU memory, and 0.327 s first-frame model latency, compared with 7.80 FPS, 24.447 GB, and 2.83 s for FlashVSR-Tiny. The code is available at https://github.com/kopperx/RelayVSR.
الكلمات المفتاحية
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Wang, X., Li, X., Lang, Z., Yao, S., Li, H., & Chen, Z. (2026). RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution. https://omanscience.com/ar/articles/relayvsr-large-small-model-collaboration-for-efficient-real-world-video-super-resolution
MLA 9
Wang, Xijun, et al. "RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution." https://omanscience.com/ar/articles/relayvsr-large-small-model-collaboration-for-efficient-real-world-video-super-resolution.
شيكاغو (المؤلف–التاريخ)
Wang, Xijun, Xin Li, Zirui Lang, Suhang Yao, Haoran Li, and Zhibo Chen. 2026. "RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution." https://omanscience.com/ar/articles/relayvsr-large-small-model-collaboration-for-efficient-real-world-video-super-resolution.
هارفارد
Wang, X., Li, X., Lang, Z., Yao, S., Li, H. and Chen, Z. (2026) 'RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution', Available at: https://omanscience.com/ar/articles/relayvsr-large-small-model-collaboration-for-efficient-real-world-video-super-resolution.
فانكوفر
Wang X, Li X, Lang Z, Yao S, Li H, Chen Z. RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution. https://omanscience.com/ar/articles/relayvsr-large-small-model-collaboration-for-efficient-real-world-video-super-resolution
IEEE
X. Wang, X. Li, Z. Lang, S. Yao, H. Li, and Z. Chen, "RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution," https://omanscience.com/ar/articles/relayvsr-large-small-model-collaboration-for-efficient-real-world-video-super-resolution.