The THU-HCSI Multi-Speaker Multi-Lingual Few-Shot Voice Cloning System for LIMMITS'24 Challenge

2024-04-25 14:02:25

Yixuan Zhou, Shuoyi Zhou, Shun Lei, Zhiyong Wu, Menglin Wu

arXiv_SD

Abstract
Abstract (translated)
URL
PDF

Abstract

This paper presents the multi-speaker multi-lingual few-shot voice cloning system developed by THU-HCSI team for LIMMITS'24 Challenge. To achieve high speaker similarity and naturalness in both mono-lingual and cross-lingual scenarios, we build the system upon YourTTS and add several enhancements. For further improving speaker similarity and speech quality, we introduce speaker-aware text encoder and flow-based decoder with Transformer blocks. In addition, we denoise the few-shot data, mix up them with pre-training data, and adopt a speaker-balanced sampling strategy to guarantee effective fine-tuning for target speakers. The official evaluations in track 1 show that our system achieves the best speaker similarity MOS of 4.25 and obtains considerable naturalness MOS of 3.97.

Abstract (translated)

本文介绍了由THU-HCSI团队为LIMMITS'24挑战开发的的多语种、多声道语音克隆系统。为了在单语种和跨语种场景下实现高说话者相似度和自然度，我们在YourTTS基础上进行了系统构建，并添加了几个增强功能。为了进一步提高说话者相似度和语音质量，我们引入了说话者感知的文本编码器和基于Transformer的流式解码器。此外，我们还对几 shot数据进行了去噪、混合处理，并采用了一种针对说话者的平衡采样策略，以确保对目标说话者的有效微调。在1号轨道的官方评估中，我们的系统实现了4.25的说话者相似度MOS和显著的自然度MOS。

URL

https://arxiv.org/abs/2404.16619

PDF

https://arxiv.org/pdf/2404.16619.pdf

The THU-HCSI Multi-Speaker Multi-Lingual Few-Shot Voice Cloning System for LIMMITS'24 Challenge

Abstract

Abstract (translated)

URL

PDF Copy

PDF