Abstract
Combination approaches for speech recognition (ASR) systems cover structured sentence-level or word-based merging techniques as well as combination of model scores during beam search. In this work, we compare model combination across popular ASR architectures. Our method leverages the complementary strengths of different models in exploring diverse portions of the search space. We rescore a joint hypothesis list of two model candidates. We then identify the best hypothesis through log-linear combination of these sequence-level scores. While model combination during first-pass recognition may yield improved performance, it introduces variability due to differing decoding methods, making direct comparison more challenging. Our two-pass method ensures consistent comparisons across all system combination results presented in this study. We evaluate model pair candidates with varying architectures and label topologies and units. Experimental results are provided for the Librispeech 960h task.
Abstract (translated)
语音识别(ASR)系统的组合方法涵盖了结构化的句子级或基于单词的合并技术,以及在束搜索过程中模型分数的结合。在这项工作中,我们比较了不同流行ASR架构中的模型组合方式。我们的方法利用了不同模型探索搜索空间中多样部分的优势。我们将重新评分两个候选模型的联合假设列表,然后通过这些序列级别得分的对数线性组合来确定最佳假设。 虽然在第一次解码过程中进行模型组合可能会提高性能,但由于不同的解码方法导致的变化性,使得直接比较更加困难。我们的两步法确保了所有系统组合结果在这项研究中的对比一致性。我们评估具有不同架构、标签拓扑和单元的模型对候选者,并提供了Librispeech 960小时任务的实验结果。 总的来说,这项工作旨在通过采用多种模型组合策略来改进语音识别系统的性能,同时提供了一种一致的方法来比较不同ASR系统的效果。
URL
https://arxiv.org/abs/2508.09880