Abstract
We present Voxtral Mini and Voxtral Small, two multimodal audio chat models. Voxtral is trained to comprehend both spoken audio and text documents, achieving state-of-the-art performance across a diverse range of audio benchmarks, while preserving strong text capabilities. Voxtral Small outperforms a number of closed-source models, while being small enough to run locally. A 32K context window enables the model to handle audio files up to 40 minutes in duration and long multi-turn conversations. We also contribute three benchmarks for evaluating speech understanding models on knowledge and trivia. Both Voxtral models are released under Apache 2.0 license.
Abstract (translated)
我们介绍了Voxtral Mini和Voxtral Small,这两款多模态音频聊天模型。Voxtral经过训练能够理解语音音频和文本文档,在各种音频基准测试中表现出色的同时,保持了强大的文本处理能力。Voxtral Small在性能上超过了多个闭源模型,并且足够小巧可以在本地运行。该模型配备了32K的上下文窗口,使其能够处理长达40分钟的音频文件以及长时间多轮对话。我们还提供了三个基准测试来评估语音理解模型在知识和趣味问答方面的表现能力。Voxtral两个版本均以Apache 2.0许可证发布。
URL
https://arxiv.org/abs/2507.13264