Abstract

We ask: given a retrieved source audio $S$ and a separate reference audio $R$, can we synthesize novel audio $Y$ out of this pair $(S,R)$ such that $Y$ remains acoustically consistent with $S$, while not persistently copying segments of $S$ or $R$? The first clause is a well-known goal in Foley audio production, and the second is a well-known issue in neural RAG when $S$ and $R$ are naively injected into neural generators. We show that both clauses can be addressed simultaneously using a method we coin relational synthesis, a variation of concatenative synthesis where target cost is replaced by a relational Gromov-like structural cost. Rather than imitating the content of $R$, relational synthesis exploits it from the "other side of the hill": it transfers the temporal structure and directed amplitude motion of $R$ to reorganize and concatenate the grains of $S$ in a novel manner that protects $S$'s acoustic information. Our experiments show that relational synthesis integrates naturally with neural RAG and produces Foley audio that performs well on metrics measuring temporal agreement, acoustic fidelity, and leakage persistence, while maintaining distribution-level quality and text alignment.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Shao, K., Kawano, A., & Dubnov, S. (2026). Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation. https://omanscience.com/en/articles/relational-synthesis-structure-mediated-concatenative-synthesis-for-foley-and-retrieval-augmented-audio-generation

MLA 9

Shao, Keren, et al. "Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation." https://omanscience.com/en/articles/relational-synthesis-structure-mediated-concatenative-synthesis-for-foley-and-retrieval-augmented-audio-generation.

Chicago (author–date)

Shao, Keren, Ayaka Kawano, and Shlomo Dubnov. 2026. "Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation." https://omanscience.com/en/articles/relational-synthesis-structure-mediated-concatenative-synthesis-for-foley-and-retrieval-augmented-audio-generation.

Harvard

Shao, K., Kawano, A. and Dubnov, S. (2026) 'Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation', Available at: https://omanscience.com/en/articles/relational-synthesis-structure-mediated-concatenative-synthesis-for-foley-and-retrieval-augmented-audio-generation.

Vancouver

Shao K, Kawano A, Dubnov S. Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation. https://omanscience.com/en/articles/relational-synthesis-structure-mediated-concatenative-synthesis-for-foley-and-retrieval-augmented-audio-generation

IEEE

K. Shao, A. Kawano, and S. Dubnov, "Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation," https://omanscience.com/en/articles/relational-synthesis-structure-mediated-concatenative-synthesis-for-foley-and-retrieval-augmented-audio-generation.