On incorporating social speaker characteristics in synthetic speech

2022-04-03 16:51:21

Sai Sirisha Rallabandi, Sebastian Möller

arXiv_SD

arXiv_SD Quantization Speech

Abstract
Abstract (translated)
URL
PDF

Abstract

In our previous work, we derived the acoustic features, that contribute to the perception of warmth and competence in synthetic speech. As an extension, in our current work, we investigate the impact of the derived vocal features in the generation of the desired characteristics. The acoustic features, spectral flux, F1 mean and F2 mean and their convex combinations were explored for the generation of higher warmth in female speech. The voiced slope, spectral flux, and their convex combinations were investigated for the generation of higher competence in female speech. We have employed a feature quantization approach in the traditional end-to-end tacotron based speech synthesis model. The listening tests have shown that the convex combination of acoustic features displays higher Mean Opinion Scores of warmth and competence when compared to that of individual features.

Abstract (translated)

URL

https://arxiv.org/abs/2204.01115

PDF

https://arxiv.org/pdf/2204.01115.pdf