Related papers: Investigating Large Language Models in Diagnosing Students' Cognitive Skills in Math Problem-solving

Investigating Large Language Models in Diagnosing Students' Cognitive Skills in Math Problem-solving

URL: http://arxiv.org/abs/2504.00843v1
Date: Tue, 01 Apr 2025 14:29:41 GMT
Title: Investigating Large Language Models in Diagnosing Students' Cognitive Skills in Math Problem-solving
Authors: Hyoungwook Jin, Yoonsu Kim, Dongyun Jung, Seungju Kim, Kiyoon Choi, Jinho Son, Juho Kim,
Abstract summary: We investigate how state-of-the-art large language models diagnose students' cognitive skills in mathematics.<n>We constructed MathCog, a novel benchmark dataset comprising 639 student responses to 110 middle school math problems.<n>Our evaluation reveals that even the state-of-the-art LLMs struggle with the task, all F1 scores below 0.5, and tend to exhibit strong false confidence for incorrect cases.
Score: 23.811625065982486
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Abstract: Mathematics learning entails mastery of both content knowledge and cognitive processing of knowing, applying, and reasoning with it. Automated math assessment primarily has focused on grading students' exhibition of content knowledge by finding textual evidence, such as specific numbers, formulas, and statements. Recent advancements in problem-solving, image recognition, and reasoning capabilities of large language models (LLMs) show promise for nuanced evaluation of students' cognitive skills. Diagnosing cognitive skills needs to infer students' thinking processes beyond textual evidence, which is an underexplored task in LLM-based automated assessment. In this work, we investigate how state-of-the-art LLMs diagnose students' cognitive skills in mathematics. We constructed MathCog, a novel benchmark dataset comprising 639 student responses to 110 expert-curated middle school math problems, each annotated with detailed teachers' diagnoses based on cognitive skill checklists. Using MathCog, we evaluated 16 closed and open LLMs of varying model sizes and vendors. Our evaluation reveals that even the state-of-the-art LLMs struggle with the task, all F1 scores below 0.5, and tend to exhibit strong false confidence for incorrect cases ($r_s=.617$). We also found that model size positively correlates with the diagnosis performance ($r_s=.771$). Finally, we discuss the implications of these findings, the overconfidence issue, and directions for improving automated cognitive skill diagnosis.

Related papers

Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge [0.0]
Existing studies treat memorization and self-knowledge deficits in LLMs as separate issues.<n>We utilize a novel framework to ascertain if LLMs genuinely learn reasoning patterns from training data.<n>Our analysis shows a noteworthy problem in generalization: LLMs draw confidence from memorized solutions to infer a higher self-knowledge.
arXiv Detail & Related papers (2025-06-23T18:01:16Z)
CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective [68.94793547575343]
CogMath formalizes human reasoning process into 3 stages: emphproblem comprehension, emphproblem solving, and emphsolution summarization.<n>In each dimension, we develop an emphInquiry-emphJudge-emphReference'' multi-agent system to generate inquiries that assess LLMs' mastery from this dimension.<n>An LLM is considered to truly master a problem only when excelling in all inquiries from the 9 dimensions.
arXiv Detail & Related papers (2025-06-04T22:00:52Z)
SMART: Self-Generating and Self-Validating Multi-Dimensional Assessment for LLMs' Mathematical Problem Solving [25.2860558348958]
Large Language Models have achieved remarkable results on a variety of mathematical benchmarks.<n>Common evaluation metrics, such as final answer accuracy, fail to disentangle the underlying competencies involved.<n>We introduce SMART: a Self-Generating and Self-Validating Multi-Dimensional Assessment Framework.
arXiv Detail & Related papers (2025-05-22T13:18:24Z)
A Benchmark for Math Misconceptions: Bridging Gaps in Middle School Algebra with AI-Supported Instruction [0.0]
This study introduces an evaluation benchmark for middle school algebra to be used in artificial intelligence based educational platforms. The data set comprises 55 misconceptions about algebra, common errors, and 220 diagnostic examples. Four out of five educators expressed interest in using the data set with AI to diagnose student misconceptions or train teachers.
arXiv Detail & Related papers (2024-12-04T23:10:29Z)
MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education [2.872215065231376]
This paper introduces MalAlgoQA, a dataset designed to evaluate the counterfactual reasoning capabilities of Large Language Models. At the heart of MalAlgoQA are malgorithms'' - rationales behind incorrect answer choices that represent flawed yet logically coherent reasoning paths.
arXiv Detail & Related papers (2024-07-01T03:39:13Z)
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads [74.54183505245553]
A systematic analysis of AI capabilities for joint vision and text reasoning is missing in the current scientific literature.<n>We evaluate state-of-the-art LVLMs on their mathematical and algorithmic reasoning abilities using visuo-linguistic problems from children's Olympiads.<n>Our results show that modern LVLMs do demonstrate increasingly powerful reasoning skills in solving problems for higher grades, but lack the foundations to correctly answer problems designed for younger children.
arXiv Detail & Related papers (2024-06-22T05:04:39Z)
Knowledge Tagging System on Math Questions via LLMs with Flexible Demonstration Retriever [48.5585921817745]
Large Language Models (LLMs) are used to automate the knowledge tagging task. We show the strong performance of zero- and few-shot results over math questions knowledge tagging tasks. By proposing a reinforcement learning-based demonstration retriever, we successfully exploit the great potential of different-sized LLMs.
arXiv Detail & Related papers (2024-06-19T23:30:01Z)
Metacognitive Capabilities of LLMs: An Exploration in Mathematical Problem Solving [86.04158840879727]
We develop a prompt-guided interaction procedure to get a powerful LLM to assign sensible skill labels to math questions. We then have it perform semantic clustering to obtain coarser families of skill labels. These coarse skill labels look interpretable to humans.
arXiv Detail & Related papers (2024-05-20T17:45:26Z)
Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners? [140.9751389452011]
We study the biases of large language models (LLMs) in relation to those known in children when solving arithmetic word problems. We generate a novel set of word problems for each of these tests, using a neuro-symbolic approach that enables fine-grained control over the problem features.
arXiv Detail & Related papers (2024-01-31T18:48:20Z)
Three Questions Concerning the Use of Large Language Models to Facilitate Mathematics Learning [4.376598435975689]
We discuss the challenges associated with employing large language models to enhance students' mathematical problem-solving skills. LLMs can generate the wrong reasoning processes, and also exhibit difficulty in understanding the given questions' rationales when attempting to correct students' answers.
arXiv Detail & Related papers (2023-10-20T16:05:35Z)
Evaluating Language Models for Mathematics through Interactions [116.67206980096513]
We introduce CheckMate, a prototype platform for humans to interact with and evaluate large language models (LLMs) We conduct a study with CheckMate to evaluate three language models (InstructGPT, ChatGPT, and GPT-4) as assistants in proving undergraduate-level mathematics. We derive a taxonomy of human behaviours and uncover that despite a generally positive correlation, there are notable instances of divergence between correctness and perceived helpfulness.
arXiv Detail & Related papers (2023-06-02T17:12:25Z)
Do Large Language Models Know What They Don't Know? [74.65014158544011]
Large language models (LLMs) have a wealth of knowledge that allows them to excel in various Natural Language Processing (NLP) tasks. Despite their vast knowledge, LLMs are still limited by the amount of information they can accommodate and comprehend. This study aims to evaluate LLMs' self-knowledge by assessing their ability to identify unanswerable or unknowable questions.
arXiv Detail & Related papers (2023-05-29T15:30:13Z)
Computationally Identifying Funneling and Focusing Questions in Classroom Discourse [24.279653100481863]
We propose the task of computationally detecting funneling and focusing questions in classroom discourse. We release an annotated dataset of 2,348 teacher utterances labeled for funneling and focusing questions, or neither. Our best model, a supervised RoBERTa model fine-tuned on our dataset, has a strong linear correlation of.76 with human expert labels and with positive educational outcomes.
arXiv Detail & Related papers (2022-07-08T01:28:29Z)

This list is automatically generated from the titles and abstracts of the papers in this site.