Speech Technology for Everyone: Automatic Speech Recognition for Non-Native English with Transfer Learning

2021-10-01 23:11:00

Toshiko Shibano (1), Xinyi Zhang (1), Mia Taige Li (1), Haejin Cho (1), Peter Sullivan (1), Muhammad Abdul-Mageed (1) ((1) University of British Columbia)

arXiv_CL

arXiv_CL Speech_Recognition Recognition Transfer_Learning Language_Model Zero-Shot Speech

Abstract
Abstract (translated)
URL
PDF

Abstract

To address the performance gap of English ASR models on L2 English speakers, we evaluate fine-tuning of pretrained wav2vec 2.0 models (Baevski et al., 2020; Xu et al., 2021) on L2-ARCTIC, a non-native English speech corpus (Zhao et al., 2018) under different training settings. We compare \textbf{(a)} models trained with a combination of diverse accents to ones trained with only specific accents and \textbf{(b)} results from different single-accent models. Our experiments demonstrate the promise of developing ASR models for non-native English speakers, even with small amounts of L2 training data and even without a language model. Our models also excel in the zero-shot setting where we train on multiple L2 datasets and test on a blind L2 test set.

Abstract (translated)

URL

https://arxiv.org/abs/2110.00678

PDF

https://arxiv.org/pdf/2110.00678.pdf