Abstract
Large Language Models (LLMs) are increasingly utilized in AI-driven educational instruction and assessment, particularly within mathematics education. The capability of LLMs to generate accurate answers and detailed solutions for math problem-solving tasks is foundational for ensuring reliable and precise feedback and assessment in math education practices. Our study focuses on evaluating the accuracy of four LLMs (OpenAI GPT-4o and o1, DeepSeek-V3 and DeepSeek-R1) solving three categories of math tasks, including arithmetic, algebra, and number theory, and identifies step-level reasoning errors within their solutions. Instead of relying on standard benchmarks, we intentionally build math tasks (via item models) that are challenging for LLMs and prone to errors. The accuracy of final answers and the presence of errors in individual solution steps were systematically analyzed and coded. Both single-agent and dual-agent configurations were tested. It is observed that the reasoning-enhanced OpenAI o1 model consistently achieved higher or nearly perfect accuracy across all three math task categories. Analysis of errors revealed that procedural slips were the most frequent and significantly impacted overall performance, while conceptual misunderstandings were less frequent. Deploying dual-agent configurations substantially improved overall performance. These findings offer actionable insights into enhancing LLM performance and underscore effective strategies for integrating LLMs into mathematics education, thereby advancing AI-driven instructional practices and assessment precision.
Abstract (translated)
大型语言模型(LLM)在基于人工智能的教育指导和评估中的应用日益增多,尤其是在数学教育领域。LLMs生成准确答案及详细解题步骤的能力是确保数学教学实践中可靠且精确反馈与评估的基础。我们的研究重点在于评估四种LLM(OpenAI GPT-4o 和 o1、DeepSeek-V3 及 DeepSeek-R1)在算术、代数和数论三个类别中的数学任务解决准确性,并识别其解决方案中的步骤级推理错误。 不同于依赖于标准基准测试,我们特意构建了对LLMs具有挑战性且容易出错的数学题目(通过项目模型)。最终答案的准确性和各解题步骤中错误的存在情况均被系统地分析和编码。既检验了单一代理配置也检验了双代理配置的效果。观察到的是,增强推理能力的OpenAI o1 模型在所有三个数学任务类别中的准确性始终较高或接近完美。 错误分析显示,程序性失误是最常见的且显著影响整体表现的因素,而概念误解则较少出现。采用双代理配置极大地提升了总体性能。这些发现为提高LLM性能提供了可操作的见解,并强调了有效整合LLMs于数学教育中的策略,从而推进基于AI的教学实践和评估精确度的发展。
URL
https://arxiv.org/abs/2508.09932