NICE-Beam: Neural Integrated Covariance Estimators for Time-Varying Beamformers

2021-12-08 22:48:06

Jonah Casebeer, Jacob Donley, Daniel Wong, Buye Xu, Anurag Kumar

arXiv_SD

arXiv_SD Inference Knowledge Pose Enhancement Speech

Abstract
Abstract (translated)
URL
PDF

Abstract

Estimating a time-varying spatial covariance matrix for a beamforming algorithm is a challenging task, especially for wearable devices, as the algorithm must compensate for time-varying signal statistics due to rapid pose-changes. In this paper, we propose Neural Integrated Covariance Estimators for Beamformers, NICE-Beam. NICE-Beam is a general technique for learning how to estimate time-varying spatial covariance matrices, which we apply to joint speech enhancement and dereverberation. It is based on training a neural network module to non-linearly track and leverage scene information across time. We integrate our solution into a beamforming pipeline, which enables simple training, faster than real-time inference, and a variety of test-time adaptation options. We evaluate the proposed model against a suite of baselines in scenes with both stationary and moving microphones. Our results show that the proposed method can outperform a hand-tuned estimator, despite the hand-tuned estimator using oracle source separation knowledge.

Abstract (translated)

URL

https://arxiv.org/abs/2112.04613

PDF

https://arxiv.org/pdf/2112.04613.pdf