Paper Reading AI Learner

Transient Noise Removal via Diffusion-based Speech Inpainting

2025-08-12 12:25:53
Mordehay Moradi, Sharon Gannot

Abstract

In this paper, we present PGDI, a diffusion-based speech inpainting framework for restoring missing or severely corrupted speech segments. Unlike previous methods that struggle with speaker variability or long gap lengths, PGDI can accurately reconstruct gaps of up to one second in length while preserving speaker identity, prosody, and environmental factors such as reverberation. Central to this approach is classifier guidance, specifically phoneme-level guidance, which substantially improves reconstruction fidelity. PGDI operates in a speaker-independent manner and maintains robustness even when long segments are completely masked by strong transient noise, making it well-suited for real-world applications, such as fireworks, door slams, hammer strikes, and construction noise. Through extensive experiments across diverse speakers and gap lengths, we demonstrate PGDI's superior inpainting performance and its ability to handle challenging acoustic conditions. We consider both scenarios, with and without access to the transcript during inference, showing that while the availability of text further enhances performance, the model remains effective even in its absence. For audio samples, visit: this https URL

Abstract (translated)

在这篇论文中,我们提出了PGDI(基于扩散的语音填补框架),用于恢复缺失或严重损坏的语音片段。与以往的方法不同,这些方法在处理说话人变化或多段长时间间隔时表现不佳,PGDI能够在保持说话者身份、韵律和环境因素(如混响)的同时,准确重建长达一秒的间隙。该方法的核心是分类器指导,特别是音素级别的指导,这极大地提高了重构的准确性。 PGDI以独立于说话人的方式运行,并且即使在长时间段完全被强烈瞬态噪声掩盖的情况下也能保持稳健性,使其非常适合现实世界的应用场景,例如烟火声、关门声、锤击声和施工噪音。通过针对各种说话者及不同长度间隙进行广泛的实验,我们展示了PGDI卓越的填补性能及其处理复杂声学条件的能力。 本研究还探讨了推理过程中是否有访问转录文本的不同情况,结果显示虽然拥有文本可以进一步提升模型性能,但即便在没有文本的情况下,该模型依旧非常有效。音频样本请访问:[这里插入URL链接]

URL

https://arxiv.org/abs/2508.08890

PDF

https://arxiv.org/pdf/2508.08890.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot