نسخة أولية وصول مفتوح
FlowAct-R2: Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent Planning
We present FlowAct-R2, a framework for interactive humanoid video generation that combines continuous multimodal control with proactive agent planning. Our method consists of two coupled components. First, a Streaming Multimodal Reference Diffusion Transformer adapts the pretrained Seedance 2.0 Mini reference-to-video …