Scaling to Many Languages with a Triaged Multilingual Text-Dependent and Text-Independent Speaker Verification System

2021-04-05 19:48:16

Roza Chojnacka, Jason Pelecanos, Quan Wang, Ignacio Lopez Moreno

arXiv_SD

arXiv_SD Recognition Inference Knowledge Action Speech

Abstract
Abstract (translated)
URL
PDF

Abstract

In this work we study some of the challenges associated with scaling speaker recognition systems to multiple languages. To the best of our knowledge, this is the first study of speaker verification systems at the scale of 46 languages. Training models for each of the many languages can be time and energy demanding in addition to costly. Low resource languages present additional difficulties. The problem is framed from the perspective of using a smart speaker device with interactions consisting of a wake-up keyword (text-dependent) followed by a speech query (text-independent). We examine the use of a hybrid setup consisting of multilingual text-dependent and text-independent components. Experimental evidence suggests that training on multiple languages can generalize to unseen varieties while maintaining performance on seen varieties. We also found that it can reduce computational requirements for training models by an order of magnitude. Furthermore, during model inference on English data, we observe that leveraging a triage framework can reduce the number of calls to the more computationally expensive text-independent system by 73% (and reduce latency by 60%) while maintaining an EER no worse than the text-independent setup.

Abstract (translated)

URL

https://arxiv.org/abs/2104.02125

PDF

https://arxiv.org/pdf/2104.02125.pdf