Abstract
Streaming head-avatar reenactment aims to animate a reference image according to a live driving video, requiring robust motion transfer, long-term identity stability, and low latency. Existing methods often rely on specialized identity or motion representations, which can discard useful visual information and inherit failure modes from external extractors. In addition, many recent diffusion-based reenactment methods use offline, clip-based generation, jointly processing and denoising an entire video clip before producing its output, making continuous low-latency streaming difficult. We introduce PixReenact, a pixel-conditioned streaming reenactment framework built on causal video diffusion. PixReenact conditions directly on VAE-encoded reference and driving frames, without specialized identity or motion representations. To separate reference identity from driver motion, we train with cross-identity pseudo supervision together with corrective objectives anchored to the original reference and driving inputs. Long self-rollouts reduce autoregressive drift, while state-aware dual-teacher distillation separately addresses cold-start and steady-state generation. Across three cross-identity benchmarks and a long-horizon streaming benchmark, PixReenact demonstrates robust cross-identity reenactment, particularly under challenging conditions such as extreme viewpoints, occlusions, and pronounced facial expressions, while maintaining the reference identity over long streams. A 4-NFE rolling student continuously emits four frames per update with a mean emission latency of 239 ms.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Habib, G., Samuel, D., Shimshi, O., & Ben-Ari, R. (2026). PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment. https://omanscience.com/en/articles/pixreenact-pixel-conditioned-causal-video-diffusion-for-streaming-head-avatar-reenactment
MLA 9
Habib, Gavriel, et al. "PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment." https://omanscience.com/en/articles/pixreenact-pixel-conditioned-causal-video-diffusion-for-streaming-head-avatar-reenactment.
Chicago (author–date)
Habib, Gavriel, Dvir Samuel, Or Shimshi, and Rami Ben-Ari. 2026. "PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment." https://omanscience.com/en/articles/pixreenact-pixel-conditioned-causal-video-diffusion-for-streaming-head-avatar-reenactment.
Harvard
Habib, G., Samuel, D., Shimshi, O. and Ben-Ari, R. (2026) 'PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment', Available at: https://omanscience.com/en/articles/pixreenact-pixel-conditioned-causal-video-diffusion-for-streaming-head-avatar-reenactment.
Vancouver
Habib G, Samuel D, Shimshi O, Ben-Ari R. PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment. https://omanscience.com/en/articles/pixreenact-pixel-conditioned-causal-video-diffusion-for-streaming-head-avatar-reenactment
IEEE
G. Habib, D. Samuel, O. Shimshi, and R. Ben-Ari, "PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment," https://omanscience.com/en/articles/pixreenact-pixel-conditioned-causal-video-diffusion-for-streaming-head-avatar-reenactment.