Abstract
In this paper, we present PGDI, a diffusion-based speech inpainting framework for restoring missing or severely corrupted speech segments. Unlike previous methods that struggle with speaker variability or long gap lengths, PGDI can accurately reconstruct gaps of up to one second in length while preserving speaker identity, prosody, and environmental factors such as reverberation. Central to this approach is classifier guidance, specifically phoneme-level guidance, which substantially improves reconstruction fidelity. PGDI operates in a speaker-independent manner and maintains robustness even when long segments are completely masked by strong transient noise, making it well-suited for real-world applications, such as fireworks, door slams, hammer strikes, and construction noise. Through extensive experiments across diverse speakers and gap lengths, we demonstrate PGDI's superior inpainting performance and its ability to handle challenging acoustic conditions. We consider both scenarios, with and without access to the transcript during inference, showing that while the availability of text further enhances performance, the model remains effective even in its absence. For audio samples, visit: this https URL
Abstract (translated)
在这篇论文中,我们提出了PGDI(基于扩散的语音填补框架),用于恢复缺失或严重损坏的语音片段。与以往的方法不同,这些方法在处理说话人变化或多段长时间间隔时表现不佳,PGDI能够在保持说话者身份、韵律和环境因素(如混响)的同时,准确重建长达一秒的间隙。该方法的核心是分类器指导,特别是音素级别的指导,这极大地提高了重构的准确性。 PGDI以独立于说话人的方式运行,并且即使在长时间段完全被强烈瞬态噪声掩盖的情况下也能保持稳健性,使其非常适合现实世界的应用场景,例如烟火声、关门声、锤击声和施工噪音。通过针对各种说话者及不同长度间隙进行广泛的实验,我们展示了PGDI卓越的填补性能及其处理复杂声学条件的能力。 本研究还探讨了推理过程中是否有访问转录文本的不同情况,结果显示虽然拥有文本可以进一步提升模型性能,但即便在没有文本的情况下,该模型依旧非常有效。音频样本请访问:[这里插入URL链接]
URL
https://arxiv.org/abs/2508.08890