Paper Reading AI Learner

Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions

2025-08-04 07:27:26
Farah Atif, Nursultan Askarbekuly, Kareem Darwish, Monojit Choudhury

Abstract

Despite the increasing usage of Large Language Models (LLMs) in answering questions in a variety of domains, their reliability and accuracy remain unexamined for a plethora of domains including the religious domains. In this paper, we introduce a novel benchmark FiqhQA focused on the LLM generated Islamic rulings explicitly categorized by the four major Sunni schools of thought, in both Arabic and English. Unlike prior work, which either overlooks the distinctions between religious school of thought or fails to evaluate abstention behavior, we assess LLMs not only on their accuracy but also on their ability to recognize when not to answer. Our zero-shot and abstention experiments reveal significant variation across LLMs, languages, and legal schools of thought. While GPT-4o outperforms all other models in accuracy, Gemini and Fanar demonstrate superior abstention behavior critical for minimizing confident incorrect answers. Notably, all models exhibit a performance drop in Arabic, highlighting the limitations in religious reasoning for languages other than English. To the best of our knowledge, this is the first study to benchmark the efficacy of LLMs for fine-grained Islamic school of thought specific ruling generation and to evaluate abstention for Islamic jurisprudence queries. Our findings underscore the need for task-specific evaluation and cautious deployment of LLMs in religious applications.

Abstract (translated)

尽管大型语言模型(LLMs)在回答各种领域的问题时使用越来越广泛,但它们在包括宗教领域的可靠性和准确性尚未得到充分审查。本文介绍了一种新的基准测试FiqhQA,该基准专注于根据四大逊尼派教法学派明确分类的伊斯兰法规生成,并且支持阿拉伯语和英语。 与以往的研究不同的是,这些研究要么忽略了宗教学派之间的差异,要么未能评估模型不回答问题的行为。我们在准确性的同时也评估了LLMs识别何时不应作答的能力。我们的零样本和拒绝实验显示,在不同的大型语言模型、语言以及法律学派之间存在显著的差异。尽管GPT-4o在准确度上超过了所有其他模型,但Gemini和Fanar展现了更出色的拒绝行为,这对于减少自信而错误的回答至关重要。 值得注意的是,所有模型在阿拉伯语中的表现都有所下降,这表明除了英语之外的语言在宗教推理方面的能力有限。据我们所知,这是首次对大型语言模型生成特定的伊斯兰教法学派法规的有效性进行基准测试,并评估其对伊斯兰法问题拒绝回答的行为。我们的研究结果强调了针对具体任务的评估以及在宗教应用中谨慎部署LLMs的需求。

URL

https://arxiv.org/abs/2508.08287

PDF

https://arxiv.org/pdf/2508.08287.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot