Paper Reading AI Learner

A Comparative Analysis on ASR System Combination for Attention, CTC, Factored Hybrid, and Transducer Models

2025-08-13 15:26:11
Noureldin Bayoumi, Robin Schmitt, Tina Raissi, Albert Zeyer, Ralf Schl\"uter, Hermann Ney

Abstract

Combination approaches for speech recognition (ASR) systems cover structured sentence-level or word-based merging techniques as well as combination of model scores during beam search. In this work, we compare model combination across popular ASR architectures. Our method leverages the complementary strengths of different models in exploring diverse portions of the search space. We rescore a joint hypothesis list of two model candidates. We then identify the best hypothesis through log-linear combination of these sequence-level scores. While model combination during first-pass recognition may yield improved performance, it introduces variability due to differing decoding methods, making direct comparison more challenging. Our two-pass method ensures consistent comparisons across all system combination results presented in this study. We evaluate model pair candidates with varying architectures and label topologies and units. Experimental results are provided for the Librispeech 960h task.

Abstract (translated)

语音识别(ASR)系统的组合方法涵盖了结构化的句子级或基于单词的合并技术,以及在束搜索过程中模型分数的结合。在这项工作中,我们比较了不同流行ASR架构中的模型组合方式。我们的方法利用了不同模型探索搜索空间中多样部分的优势。我们将重新评分两个候选模型的联合假设列表,然后通过这些序列级别得分的对数线性组合来确定最佳假设。 虽然在第一次解码过程中进行模型组合可能会提高性能,但由于不同的解码方法导致的变化性,使得直接比较更加困难。我们的两步法确保了所有系统组合结果在这项研究中的对比一致性。我们评估具有不同架构、标签拓扑和单元的模型对候选者,并提供了Librispeech 960小时任务的实验结果。 总的来说,这项工作旨在通过采用多种模型组合策略来改进语音识别系统的性能,同时提供了一种一致的方法来比较不同ASR系统的效果。

URL

https://arxiv.org/abs/2508.09880

PDF

https://arxiv.org/pdf/2508.09880.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot