الملخص
Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise generative fidelity. We propose ReCaVSR, a Wan2.2-based, one-step framework for streaming VSR that builds on two observations: recycled SR latents retain local temporal context, reducing the need for full historical Key-Value (KV) caches; and individual transformer layers benefit from distinct temporal scopes. ReCaVSR combines three complementary designs: (i) layer-wise cache routing with recycled SR latents: each DiT layer learns its KV-cache temporal scope under a cache budget and exports a static inference schedule, while recycled SR latents propagate local context by conditioning each new block on the model's own preceding predictions. (ii) Multi-Scope Query (MSQ) Discriminator: a compositional discriminator combining global, spatial-window, and temporal-tube feedback for holistic realism, local texture generation, and temporal stability. (iii) LR-conditioned adaptation of FlashDecoder: a VAE decoder that incorporates LR observations for efficient latent decoding. ReCaVSR enables streaming VSR without iterative sampling or full historical KV-cache materialization. Experiments on synthetic and real-world VSR benchmarks show better perceptual quality, temporal consistency, and streaming efficiency than representative VSR baselines. At $1080{\times}1920$ output resolution on a single NVIDIA A100-80GB, ReCaVSR achieves 21.20 FPS with 15.16 GB peak allocated GPU memory, running 2.72$\times$ faster while using 38.0\% less peak allocated memory than FlashVSR Tiny. The code is available at https://github.com/kopperx/ReCaVSR.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Wang, X., Li, X., Yao, S., Lang, Z., Li, B., & Chen, Z. (2026). ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing. https://omanscience.com/ar/articles/recavsr-one-step-streaming-diffusion-video-super-resolution-with-recycled-latents-and-learned-cache-routing
MLA 9
Wang, Xijun, et al. "ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing." https://omanscience.com/ar/articles/recavsr-one-step-streaming-diffusion-video-super-resolution-with-recycled-latents-and-learned-cache-routing.
شيكاغو (المؤلف–التاريخ)
Wang, Xijun, Xin Li, Suhang Yao, Zirui Lang, Bingchen Li, and Zhibo Chen. 2026. "ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing." https://omanscience.com/ar/articles/recavsr-one-step-streaming-diffusion-video-super-resolution-with-recycled-latents-and-learned-cache-routing.
هارفارد
Wang, X., Li, X., Yao, S., Lang, Z., Li, B. and Chen, Z. (2026) 'ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing', Available at: https://omanscience.com/ar/articles/recavsr-one-step-streaming-diffusion-video-super-resolution-with-recycled-latents-and-learned-cache-routing.
فانكوفر
Wang X, Li X, Yao S, Lang Z, Li B, Chen Z. ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing. https://omanscience.com/ar/articles/recavsr-one-step-streaming-diffusion-video-super-resolution-with-recycled-latents-and-learned-cache-routing
IEEE
X. Wang, X. Li, S. Yao, Z. Lang, B. Li, and Z. Chen, "ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing," https://omanscience.com/ar/articles/recavsr-one-step-streaming-diffusion-video-super-resolution-with-recycled-latents-and-learned-cache-routing.