Paper Reading AI Learner

MILU: A Multi-task Indic Language Understanding Benchmark

2024-11-04 19:17:17
Sshubam Verma, Mohammed Safi Ur Rahman Khan, Vishwajeet Kumar, Rudra Murthy, Jaydeep Sen

Abstract

Evaluating Large Language Models (LLMs) in low-resource and linguistically diverse languages remains a significant challenge in NLP, particularly for languages using non-Latin scripts like those spoken in India. Existing benchmarks predominantly focus on English, leaving substantial gaps in assessing LLM capabilities in these languages. We introduce MILU, a Multi task Indic Language Understanding Benchmark, a comprehensive evaluation benchmark designed to address this gap. MILU spans 8 domains and 42 subjects across 11 Indic languages, reflecting both general and culturally specific knowledge. With an India-centric design, incorporates material from regional and state-level examinations, covering topics such as local history, arts, festivals, and laws, alongside standard subjects like science and mathematics. We evaluate over 42 LLMs, and find that current LLMs struggle with MILU, with GPT-4o achieving the highest average accuracy at 72 percent. Open multilingual models outperform language-specific fine-tuned models, which perform only slightly better than random baselines. Models also perform better in high resource languages as compared to low resource ones. Domain-wise analysis indicates that models perform poorly in culturally relevant areas like Arts and Humanities, Law and Governance compared to general fields like STEM. To the best of our knowledge, MILU is the first of its kind benchmark focused on Indic languages, serving as a crucial step towards comprehensive cultural evaluation. All code, benchmarks, and artifacts will be made publicly available to foster open research.

Abstract (translated)

评估大型语言模型(LLMs)在资源有限和语言多样化的环境中,特别是在使用非拉丁字母的印度等地区的语言中,仍然是自然语言处理(NLP)领域的一个重大挑战。现有的基准测试主要集中在英语上,这导致了对这些语言中的LLM能力评估的巨大空白。我们引入了MILU(多任务印地语理解基准),这是一个旨在解决这一问题的全面评测基准。MILU涵盖了11种印度语言中的8个领域和42个主题,反映了普遍知识以及文化特定的知识。该设计以印度为中心,纳入了来自地区及州级考试的内容,涵盖地方历史、艺术、节日和法律等主题,同时也包括科学和数学等标准科目。我们评估了超过42个LLM,并发现当前的LLMs在MILU上的表现不尽如人意,其中GPT-4o达到了最高的平均准确率72%。开放多语言模型的表现优于特定语言微调的模型,后者仅略高于随机基线水平。模型在资源丰富的语言中的表现优于资源匮乏的语言。按领域分析表明,与STEM等通用领域相比,模型在艺术和人文、法律和治理等领域这样的文化相关领域的表现较差。据我们所知,MILU是首个专注于印度语言的此类基准测试,为全面的文化评估迈出了重要一步。所有代码、基准和资源都将公开发布,以促进开放研究。

URL

https://arxiv.org/abs/2411.02538

PDF

https://arxiv.org/pdf/2411.02538.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot