Abstract

Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance on a computationally intensive content encoder remains a bottleneck. We therefore propose MeanVoiceFlow2, a framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. The model is trained through conversion distillation using MeanVoiceFlow and the reconstruction of real data. We further incorporate diffusion-GAN training with sample mixing and teacher-guided conditioning augmentation to enhance realism and disentanglement. Experiments on zero-shot VC showed that MeanVoiceFlow2 achieved higher perceptual quality and approximately $9\times$ faster inference than MeanVoiceFlow while maintaining comparable speaker similarity. Audio samples are available at https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/meanvoiceflow2/.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Kaneko, T., Kameoka, H., Tanaka, K., & Kondo, Y. (2026). MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion. https://omanscience.com/en/articles/meanvoiceflow2-joint-optimization-of-mean-flow-and-content-encoder-for-fast-one-step-zero-shot-voice-conversion

MLA 9

Kaneko, Takuhiro, et al. "MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion." https://omanscience.com/en/articles/meanvoiceflow2-joint-optimization-of-mean-flow-and-content-encoder-for-fast-one-step-zero-shot-voice-conversion.

Chicago (author–date)

Kaneko, Takuhiro, Hirokazu Kameoka, Kou Tanaka, and Yuto Kondo. 2026. "MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion." https://omanscience.com/en/articles/meanvoiceflow2-joint-optimization-of-mean-flow-and-content-encoder-for-fast-one-step-zero-shot-voice-conversion.

Harvard

Kaneko, T., Kameoka, H., Tanaka, K. and Kondo, Y. (2026) 'MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion', Available at: https://omanscience.com/en/articles/meanvoiceflow2-joint-optimization-of-mean-flow-and-content-encoder-for-fast-one-step-zero-shot-voice-conversion.

Vancouver

Kaneko T, Kameoka H, Tanaka K, Kondo Y. MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion. https://omanscience.com/en/articles/meanvoiceflow2-joint-optimization-of-mean-flow-and-content-encoder-for-fast-one-step-zero-shot-voice-conversion

IEEE

T. Kaneko, H. Kameoka, K. Tanaka, and Y. Kondo, "MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion," https://omanscience.com/en/articles/meanvoiceflow2-joint-optimization-of-mean-flow-and-content-encoder-for-fast-one-step-zero-shot-voice-conversion.