Abstract
Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed track. This limits source-level control and differs from practical audiovisual workflows, where speech, music, sound effects, and ambient sounds are represented as separate editable tracks. We introduce Soundwich, a training-free framework that transforms a frozen joint audio-video flow-matching model into a generator of multiple synchronized, independently editable audio stems coupled to a shared video. Soundwich generates separate audio stems with explicit control over their temporal activity. To keep separately generated sounds coherent, we introduce a shared scene representation that communicates global audiovisual context across stems while preserving their source-level separation. We further route cross-modal interactions between each audio stem and its corresponding visual source, improving audiovisual consistency. The resulting stems remain synchronized with the video and can be independently retimed, muted, replaced, or remixed. Experiments and human evaluations show improved temporal control, source separation, and naturalness, while enabling flexible source-level editing within coherent audiovisual generation. Code is available at https://github.com/CodyNing/Soundwich.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Ning, Z., Razlighi, A. N., Polaczek, S., Cohen-Or, D., & Mahdavi-Amiri, A. (2026). Soundwich: Video Generation with Layered and Controllable Audio. https://omanscience.com/en/articles/soundwich-video-generation-with-layered-and-controllable-audio
MLA 9
Ning, Zhuo, et al. "Soundwich: Video Generation with Layered and Controllable Audio." https://omanscience.com/en/articles/soundwich-video-generation-with-layered-and-controllable-audio.
Chicago (author–date)
Ning, Zhuo, AmirHossein Naghi Razlighi, Sagi Polaczek, Daniel Cohen-Or, and Ali Mahdavi-Amiri. 2026. "Soundwich: Video Generation with Layered and Controllable Audio." https://omanscience.com/en/articles/soundwich-video-generation-with-layered-and-controllable-audio.
Harvard
Ning, Z., Razlighi, A. N., Polaczek, S., Cohen-Or, D. and Mahdavi-Amiri, A. (2026) 'Soundwich: Video Generation with Layered and Controllable Audio', Available at: https://omanscience.com/en/articles/soundwich-video-generation-with-layered-and-controllable-audio.
Vancouver
Ning Z, Razlighi AN, Polaczek S, Cohen-Or D, Mahdavi-Amiri A. Soundwich: Video Generation with Layered and Controllable Audio. https://omanscience.com/en/articles/soundwich-video-generation-with-layered-and-controllable-audio
IEEE
Z. Ning, A. N. Razlighi, S. Polaczek, D. Cohen-Or, and A. Mahdavi-Amiri, "Soundwich: Video Generation with Layered and Controllable Audio," https://omanscience.com/en/articles/soundwich-video-generation-with-layered-and-controllable-audio.