Abstract
Current multi-channel speech enhancement systems mainly adopt single-output architecture, which face significant challenges in preserving spatio-temporal signal integrity during multiple-input multiple-output (MIMO) processing. To address this limitation, we propose a novel neural network, termed WTFormer, for MIMO speech enhancement that leverages the multi-resolution characteristics of wavelet transform and multi-dimensional collaborative attention to effectively capture globally distributed spatial features, while using Conformer for time-frequency modeling. A multi task loss strategy accompanying MUSIC algorithm is further proposed for optimization training to protect spatial information to the greatest extent. Experimental results on the LibriSpeech dataset show that WTFormer can achieve comparable denoising performance to advanced systems while preserving more spatial information with only 0.98M parameters.
Abstract (translated)
当前的多通道语音增强系统主要采用单输出架构,在进行多输入多输出(MIMO)处理时面临保持时空信号完整性的重大挑战。为了解决这一限制,我们提出了一种新型神经网络WTFormer,用于MIMO语音增强。该网络利用小波变换的多分辨率特性和多维度协作注意机制来有效捕捉全局分布的空间特征,并使用Conformer进行时间-频率建模。此外,为了保护空间信息,我们还提出了一个结合MUSIC算法的多任务损失策略来进行优化训练。 在LibriSpeech数据集上的实验结果表明,WTFormer可以在仅用0.98M参数的情况下达到与先进系统相当的去噪性能,并且保留更多的空间信息。
URL
https://arxiv.org/abs/2506.22001