Evaluating GPT- and Reasoning-based Large Language Models on Physics Olympiad Problems: Surpassing Human Performance and Implications for Educational Assessment
arXiv:2505.09438 · doi:10.1103/6fmx-bsnl
Abstract
Large language models (LLMs) are now widely accessible, reaching learners at all educational levels. This development has raised concerns that their use may circumvent essential learning processes and compromise the integrity of established assessment formats. In physics education, where problem solving plays a central role in instruction and assessment, it is therefore essential to understand the physics-specific problem-solving capabilities of LLMs. Such understanding is key to informing responsible and pedagogically sound approaches to integrating LLMs into instruction and assessment. This study therefore compares the problem-solving performance of a general-purpose LLM (GPT-4o, using varying prompting techniques) and a reasoning-optimized model (o1-preview) with that of participants of the German Physics Olympiad, based on a set of well-defined Olympiad problems. In addition to evaluating the correctness of the generated solutions, the study analyzes characteristic strengths and limitations of LLM-generated solutions. The findings of this study indicate that both tested LLMs (GPT-4o and o1-preview) demonstrate advanced problem-solving capabilities on Olympiad-type physics problems, on average outperforming the human participants. Prompting techniques had little effect on GPT-4o's performance, while o1-preview almost consistently outperformed both GPT-4o and the human benchmark. Based on these findings, the study discusses implications for the design of summative and formative assessment in physics education, including how to uphold assessment integrity and support students in critically engaging with LLMs.
References in corpus (15)
- Beware of Metacognitive Laziness: Effects of Generative Artificial Intelligence on Learning Motivation, Processes, and Performance
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models
- The Debate Over Understanding in AI's Large Language Models
- Identifying 21st century STEM competencies using workplace data
- Could an Artificial-Intelligence agent pass an introductory physics course?
- Evaluating Large Language Models on a Highly-specialized Topic, Radiation Oncology Physics
- How understanding large language models can inform the use of ChatGPT in physics education
- When physical intuition fails
- The Impact of AI in Physics Education: A Comprehensive Review from GCSE to University Levels
- Performance of ChatGPT on the Test of Understanding Graphs in Kinematics
- Thematic Analysis of 18 Years of PERC Proceedings using Natural Language Processing
- Using Large Language Models to Assign Partial Credit to Students' Explanations of Problem-Solving Process: Grade at Human Level Accuracy with Grading Confidence Index and Personalized Student-facing Feedback
- Multilingual Performance of a Multimodal Artificial Intelligence System on Multisubject Physics Concept Inventories
- Cheat sites and artificial intelligence usage in online introductory physics courses: what is the extent and what effect does it have on assessments?
- Innovative approaches to high school physics competitions: Harnessing the power of AI and open science