Self-supervised Learning of Visual Speech Features with Audiovisual Speech Enhancement

2020-04-25 01:00:03

Zakaria Aldeneh, Anushree Prasanna Kumar, Barry-John Theobald, Erik Marchi, Sachin Kajarekar, Devang Naik, Ahmed Hussen Abdelaziz

arXiv_CV

arXiv_CV Embedding Activity Self-Supervised Enhancement Speech

Abstract
Abstract (translated)
URL
PDF

Abstract

We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target speech signal. We show that visual features provide not only high-level information about speech activity, i.e. speech vs. no speech, but also fine-grained visual information about the place of articulation. An interesting byproduct of this finding is that the learned visual embeddings can be used as features for other visual speech applications. We demonstrate the effectiveness of the learned visual representations for classifying visemes (the visual analogy to phonemes). Our results provide insight into important aspects of audiovisual speech enhancement and demonstrate how such models can be used for self-supervision tasks for visual speech applications.

Abstract (translated)

URL

https://arxiv.org/abs/2004.12031

PDF

https://arxiv.org/pdf/2004.12031.pdf

Self-supervised Learning of Visual Speech Features with Audiovisual Speech Enhancement

Abstract

Abstract (translated)

URL

PDF Copy

PDF