Abstract
Evaluating Large Language Models (LLMs) in low-resource and linguistically diverse languages remains a significant challenge in NLP, particularly for languages using non-Latin scripts like those spoken in India. Existing benchmarks predominantly focus on English, leaving substantial gaps in assessing LLM capabilities in these languages. We introduce MILU, a Multi task Indic Language Understanding Benchmark, a comprehensive evaluation benchmark designed to address this gap. MILU spans 8 domains and 42 subjects across 11 Indic languages, reflecting both general and culturally specific knowledge. With an India-centric design, incorporates material from regional and state-level examinations, covering topics such as local history, arts, festivals, and laws, alongside standard subjects like science and mathematics. We evaluate over 42 LLMs, and find that current LLMs struggle with MILU, with GPT-4o achieving the highest average accuracy at 72 percent. Open multilingual models outperform language-specific fine-tuned models, which perform only slightly better than random baselines. Models also perform better in high resource languages as compared to low resource ones. Domain-wise analysis indicates that models perform poorly in culturally relevant areas like Arts and Humanities, Law and Governance compared to general fields like STEM. To the best of our knowledge, MILU is the first of its kind benchmark focused on Indic languages, serving as a crucial step towards comprehensive cultural evaluation. All code, benchmarks, and artifacts will be made publicly available to foster open research.
Abstract (translated)
评估大型语言模型(LLMs)在资源有限和语言多样化的环境中,特别是在使用非拉丁字母的印度等地区的语言中,仍然是自然语言处理(NLP)领域的一个重大挑战。现有的基准测试主要集中在英语上,这导致了对这些语言中的LLM能力评估的巨大空白。我们引入了MILU(多任务印地语理解基准),这是一个旨在解决这一问题的全面评测基准。MILU涵盖了11种印度语言中的8个领域和42个主题,反映了普遍知识以及文化特定的知识。该设计以印度为中心,纳入了来自地区及州级考试的内容,涵盖地方历史、艺术、节日和法律等主题,同时也包括科学和数学等标准科目。我们评估了超过42个LLM,并发现当前的LLMs在MILU上的表现不尽如人意,其中GPT-4o达到了最高的平均准确率72%。开放多语言模型的表现优于特定语言微调的模型,后者仅略高于随机基线水平。模型在资源丰富的语言中的表现优于资源匮乏的语言。按领域分析表明,与STEM等通用领域相比,模型在艺术和人文、法律和治理等领域这样的文化相关领域的表现较差。据我们所知,MILU是首个专注于印度语言的此类基准测试,为全面的文化评估迈出了重要一步。所有代码、基准和资源都将公开发布,以促进开放研究。
URL
https://arxiv.org/abs/2411.02538