Preprint Open access
Learning to Watermark Speech Synthesis Against Model-Driven Reconstruction
Modern TTS systems increasingly generate synthetic speech at scale for diverse users. This setting calls for content-level provenance that can verify the origin of released speech and attribute it to the requesting user, which generative watermarking can support by embedding multi-bit identifiers directly into synthesi …