الملخص

Streaming head-avatar reenactment aims to animate a reference image according to a live driving video, requiring robust motion transfer, long-term identity stability, and low latency. Existing methods often rely on specialized identity or motion representations, which can discard useful visual information and inherit failure modes from external extractors. In addition, many recent diffusion-based reenactment methods use offline, clip-based generation, jointly processing and denoising an entire video clip before producing its output, making continuous low-latency streaming difficult. We introduce PixReenact, a pixel-conditioned streaming reenactment framework built on causal video diffusion. PixReenact conditions directly on VAE-encoded reference and driving frames, without specialized identity or motion representations. To separate reference identity from driver motion, we train with cross-identity pseudo supervision together with corrective objectives anchored to the original reference and driving inputs. Long self-rollouts reduce autoregressive drift, while state-aware dual-teacher distillation separately addresses cold-start and steady-state generation. Across three cross-identity benchmarks and a long-horizon streaming benchmark, PixReenact demonstrates robust cross-identity reenactment, particularly under challenging conditions such as extreme viewpoints, occlusions, and pronounced facial expressions, while maintaining the reference identity over long streams. A 4-NFE rolling student continuously emits four frames per update with a mean emission latency of 239 ms.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Habib, G., Samuel, D., Shimshi, O., & Ben-Ari, R. (2026). PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment. https://omanscience.com/ar/articles/pixreenact-pixel-conditioned-causal-video-diffusion-for-streaming-head-avatar-reenactment

MLA 9

Habib, Gavriel, et al. "PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment." https://omanscience.com/ar/articles/pixreenact-pixel-conditioned-causal-video-diffusion-for-streaming-head-avatar-reenactment.

شيكاغو (المؤلف–التاريخ)

Habib, Gavriel, Dvir Samuel, Or Shimshi, and Rami Ben-Ari. 2026. "PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment." https://omanscience.com/ar/articles/pixreenact-pixel-conditioned-causal-video-diffusion-for-streaming-head-avatar-reenactment.

هارفارد

Habib, G., Samuel, D., Shimshi, O. and Ben-Ari, R. (2026) 'PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment', Available at: https://omanscience.com/ar/articles/pixreenact-pixel-conditioned-causal-video-diffusion-for-streaming-head-avatar-reenactment.

فانكوفر

Habib G, Samuel D, Shimshi O, Ben-Ari R. PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment. https://omanscience.com/ar/articles/pixreenact-pixel-conditioned-causal-video-diffusion-for-streaming-head-avatar-reenactment

IEEE

G. Habib, D. Samuel, O. Shimshi, and R. Ben-Ari, "PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment," https://omanscience.com/ar/articles/pixreenact-pixel-conditioned-causal-video-diffusion-for-streaming-head-avatar-reenactment.