Perceptual Contrast Stretching on Target Feature for Speech Enhancement

2022-03-31 16:24:51

Rong Chao, Cheng Yu, Szu-Wei Fu, Xugang Lu, Yu Tsao

arXiv_SD

Abstract
Abstract (translated)
URL
PDF

Abstract

Speech enhancement (SE) performance has improved considerably since the use of deep learning (DL) models as a base function. In this study, we propose a perceptual contrast stretching (PCS) approach to further improve SE performance. PCS is derived based on the critical band importance function and applied to modify the targets of the SE model. Specifically, PCS stretches the contract of target features according to perceptual importance, thereby improving the overall SE performance. Compared to post-processing based implementations, incorporating PCS into the training phase preserves performance and reduces online computation. It is also worth noting that PCS can be suitably combined with different SE model architectures and training criteria. Meanwhile, PCS does not affect the causality or convergence of the SE model training. Experimental results on the VoiceBank-DEMAND dataset showed that the proposed method can achieve state-of-the-art performance on both causal (PESQ=3.07) and non-causal (PESQ=3.35) SE tasks.

Abstract (translated)

URL

https://arxiv.org/abs/2203.17152

PDF

https://arxiv.org/pdf/2203.17152.pdf