Abstract
We ask: given a retrieved source audio $S$ and a separate reference audio $R$, can we synthesize novel audio $Y$ out of this pair $(S,R)$ such that $Y$ remains acoustically consistent with $S$, while not persistently copying segments of $S$ or $R$? The first clause is a well-known goal in Foley audio production, and the second is a well-known issue in neural RAG when $S$ and $R$ are naively injected into neural generators. We show that both clauses can be addressed simultaneously using a method we coin relational synthesis, a variation of concatenative synthesis where target cost is replaced by a relational Gromov-like structural cost. Rather than imitating the content of $R$, relational synthesis exploits it from the "other side of the hill": it transfers the temporal structure and directed amplitude motion of $R$ to reorganize and concatenate the grains of $S$ in a novel manner that protects $S$'s acoustic information. Our experiments show that relational synthesis integrates naturally with neural RAG and produces Foley audio that performs well on metrics measuring temporal agreement, acoustic fidelity, and leakage persistence, while maintaining distribution-level quality and text alignment.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Shao, K., Kawano, A., & Dubnov, S. (2026). Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation. https://omanscience.com/en/articles/relational-synthesis-structure-mediated-concatenative-synthesis-for-foley-and-retrieval-augmented-audio-generation
MLA 9
Shao, Keren, et al. "Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation." https://omanscience.com/en/articles/relational-synthesis-structure-mediated-concatenative-synthesis-for-foley-and-retrieval-augmented-audio-generation.
Chicago (author–date)
Shao, Keren, Ayaka Kawano, and Shlomo Dubnov. 2026. "Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation." https://omanscience.com/en/articles/relational-synthesis-structure-mediated-concatenative-synthesis-for-foley-and-retrieval-augmented-audio-generation.
Harvard
Shao, K., Kawano, A. and Dubnov, S. (2026) 'Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation', Available at: https://omanscience.com/en/articles/relational-synthesis-structure-mediated-concatenative-synthesis-for-foley-and-retrieval-augmented-audio-generation.
Vancouver
Shao K, Kawano A, Dubnov S. Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation. https://omanscience.com/en/articles/relational-synthesis-structure-mediated-concatenative-synthesis-for-foley-and-retrieval-augmented-audio-generation
IEEE
K. Shao, A. Kawano, and S. Dubnov, "Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation," https://omanscience.com/en/articles/relational-synthesis-structure-mediated-concatenative-synthesis-for-foley-and-retrieval-augmented-audio-generation.