Paper Reading AI Learner

Mathematical Computation and Reasoning Errors by Large Language Models

2025-08-13 16:33:02
Liang Zhang, Edith Aurora Graf

Abstract

Large Language Models (LLMs) are increasingly utilized in AI-driven educational instruction and assessment, particularly within mathematics education. The capability of LLMs to generate accurate answers and detailed solutions for math problem-solving tasks is foundational for ensuring reliable and precise feedback and assessment in math education practices. Our study focuses on evaluating the accuracy of four LLMs (OpenAI GPT-4o and o1, DeepSeek-V3 and DeepSeek-R1) solving three categories of math tasks, including arithmetic, algebra, and number theory, and identifies step-level reasoning errors within their solutions. Instead of relying on standard benchmarks, we intentionally build math tasks (via item models) that are challenging for LLMs and prone to errors. The accuracy of final answers and the presence of errors in individual solution steps were systematically analyzed and coded. Both single-agent and dual-agent configurations were tested. It is observed that the reasoning-enhanced OpenAI o1 model consistently achieved higher or nearly perfect accuracy across all three math task categories. Analysis of errors revealed that procedural slips were the most frequent and significantly impacted overall performance, while conceptual misunderstandings were less frequent. Deploying dual-agent configurations substantially improved overall performance. These findings offer actionable insights into enhancing LLM performance and underscore effective strategies for integrating LLMs into mathematics education, thereby advancing AI-driven instructional practices and assessment precision.

Abstract (translated)

大型语言模型(LLM)在基于人工智能的教育指导和评估中的应用日益增多,尤其是在数学教育领域。LLMs生成准确答案及详细解题步骤的能力是确保数学教学实践中可靠且精确反馈与评估的基础。我们的研究重点在于评估四种LLM(OpenAI GPT-4o 和 o1、DeepSeek-V3 及 DeepSeek-R1)在算术、代数和数论三个类别中的数学任务解决准确性,并识别其解决方案中的步骤级推理错误。 不同于依赖于标准基准测试,我们特意构建了对LLMs具有挑战性且容易出错的数学题目(通过项目模型)。最终答案的准确性和各解题步骤中错误的存在情况均被系统地分析和编码。既检验了单一代理配置也检验了双代理配置的效果。观察到的是,增强推理能力的OpenAI o1 模型在所有三个数学任务类别中的准确性始终较高或接近完美。 错误分析显示,程序性失误是最常见的且显著影响整体表现的因素,而概念误解则较少出现。采用双代理配置极大地提升了总体性能。这些发现为提高LLM性能提供了可操作的见解,并强调了有效整合LLMs于数学教育中的策略,从而推进基于AI的教学实践和评估精确度的发展。

URL

https://arxiv.org/abs/2508.09932

PDF

https://arxiv.org/pdf/2508.09932.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot