Abstract
Despite the increasing usage of Large Language Models (LLMs) in answering questions in a variety of domains, their reliability and accuracy remain unexamined for a plethora of domains including the religious domains. In this paper, we introduce a novel benchmark FiqhQA focused on the LLM generated Islamic rulings explicitly categorized by the four major Sunni schools of thought, in both Arabic and English. Unlike prior work, which either overlooks the distinctions between religious school of thought or fails to evaluate abstention behavior, we assess LLMs not only on their accuracy but also on their ability to recognize when not to answer. Our zero-shot and abstention experiments reveal significant variation across LLMs, languages, and legal schools of thought. While GPT-4o outperforms all other models in accuracy, Gemini and Fanar demonstrate superior abstention behavior critical for minimizing confident incorrect answers. Notably, all models exhibit a performance drop in Arabic, highlighting the limitations in religious reasoning for languages other than English. To the best of our knowledge, this is the first study to benchmark the efficacy of LLMs for fine-grained Islamic school of thought specific ruling generation and to evaluate abstention for Islamic jurisprudence queries. Our findings underscore the need for task-specific evaluation and cautious deployment of LLMs in religious applications.
Abstract (translated)
尽管大型语言模型(LLMs)在回答各种领域的问题时使用越来越广泛,但它们在包括宗教领域的可靠性和准确性尚未得到充分审查。本文介绍了一种新的基准测试FiqhQA,该基准专注于根据四大逊尼派教法学派明确分类的伊斯兰法规生成,并且支持阿拉伯语和英语。 与以往的研究不同的是,这些研究要么忽略了宗教学派之间的差异,要么未能评估模型不回答问题的行为。我们在准确性的同时也评估了LLMs识别何时不应作答的能力。我们的零样本和拒绝实验显示,在不同的大型语言模型、语言以及法律学派之间存在显著的差异。尽管GPT-4o在准确度上超过了所有其他模型,但Gemini和Fanar展现了更出色的拒绝行为,这对于减少自信而错误的回答至关重要。 值得注意的是,所有模型在阿拉伯语中的表现都有所下降,这表明除了英语之外的语言在宗教推理方面的能力有限。据我们所知,这是首次对大型语言模型生成特定的伊斯兰教法学派法规的有效性进行基准测试,并评估其对伊斯兰法问题拒绝回答的行为。我们的研究结果强调了针对具体任务的评估以及在宗教应用中谨慎部署LLMs的需求。
URL
https://arxiv.org/abs/2508.08287