Abstract
Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource languages, reaching state-of-the-art performance in various tasks. However, their applicability is still less explored in low-resource settings. This work investigates the use of Speech LLMs for low-resource Automatic Speech Recognition using the SLAM-ASR framework, where a trainable lightweight projector connects a speech encoder and a LLM. Firstly, we assess training data volume requirements to match Whisper-only performance, re-emphasizing the challenges of limited data. Secondly, we show that leveraging mono- or multilingual projectors pretrained on high-resource languages reduces the impact of data scarcity, especially with small training sets. Using multilingual LLMs (EuroLLM, Salamandra) with whisper-large-v3-turbo, we evaluate performance on several public benchmarks, providing insights for future research on optimizing Speech LLMs for low-resource languages and multilinguality.
Abstract (translated)
大型语言模型(LLMs)在处理高资源语言的口语输入方面已显示出潜力,并且在各种任务中达到了最先进的性能。然而,它们在低资源环境中的适用性研究仍然较少。本项工作探讨了使用SLAM-ASR框架下的语音LLM进行低资源自动语音识别的应用情况,在该框架下,一个可训练的轻量级投影器连接了一个语音编码器和一个大型语言模型。首先,我们评估了为了达到仅使用Whisper时的性能所需的数据量,重新强调了数据有限带来的挑战。其次,我们证明了利用在高资源语言上预训练的单语或多语种投影器可以减少数据稀缺性的影响,特别是在小规模训练集的情况下尤为明显。通过将多语言LLM(EuroLLM、Salamandra)与whisper-large-v3-turbo结合使用,并在多个公开基准测试中评估其性能,为未来优化低资源语言和多语种环境下语音LLMs的研究提供了见解。
URL
https://arxiv.org/abs/2508.05149