نسخة أولية وصول مفتوح
VINCIE-NExT: Unlocking Video Editing from Images via In-Context Modeling
Building a capable video editor remains significantly harder than a video generator: editing requires (source, instruction, edited) triplets that are prohibitively expensive to annotate and difficult to synthesize at scale, whereas image editing has already reached maturity with millions of such pairs readily available …