The Impact of AI in Physics Education: A Comprehensive Review from GCSE to University Levels
arXiv:2309.05163 · doi:10.1088/1361-6552/ad1fa2
Abstract
With the rapid evolution of Artificial Intelligence (AI), its potential implications for higher education have become a focal point of interest. This study delves into the capabilities of AI in Physics Education and offers actionable AI policy recommendations. Using a Large Language Model (LLM), we assessed its ability to answer 1337 Physics exam questions spanning GCSE, A-Level, and Introductory University curricula. We employed various AI prompting techniques: Zero Shot, In Context Learning, and Confirmatory Checking, which merges Chain of Thought reasoning with Reflection. The AI's proficiency varied across academic levels: it scored an average of 83.4% on GCSE, 63.8% on A-Level, and 37.4% on university-level questions, with an overall average of 59.9% using the most effective prompting technique. In a separate test, the LLM's accuracy on 5000 mathematical operations was found to decrease as the number of digits increased. Furthermore, when evaluated as a marking tool, the LLM's concordance with human markers averaged at 50.8%, with notable inaccuracies in marking straightforward questions, like multiple-choice. Given these results, our recommendations underscore caution: while current LLMs can consistently perform well on Physics questions at earlier educational stages, their efficacy diminishes with advanced content and complex calculations. LLM outputs often showcase novel methods not in the syllabus, excessive verbosity, and miscalculations in basic arithmetic. This suggests that at university, there's no substantial threat from LLMs for non-invigilated Physics questions. However, given the LLMs' considerable proficiency in writing Physics essays and coding abilities, non-invigilated examinations of these skills in Physics are highly vulnerable to automated completion by LLMs. This vulnerability also extends to Physics questions pitched at lower academic levels.
22 pages, 10 Figures, 2 Tables
References in corpus (13)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Evaluating Large Language Models Trained on Code
- Large Language Models are Zero-Shot Reasoners
- Gemini: A Family of Highly Capable Multimodal Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Mathematical Capabilities of ChatGPT
- A Survey on In-context Learning
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis
- The Death of the Short-Form Physics Essay in the Coming AI Revolution
- Is ChatGPT the Ultimate Programming Assistant -- How far is it?
- How do physics students evaluate artificial intelligence responses on comprehension questions? A study on the perceived scientific accuracy and linguistic quality
- On the Detectability of ChatGPT Content: Benchmarking, Methodology, and Evaluation through the Lens of Academic Writing
- Exploring Durham University Physics exams with Large Language Models
Cited by in corpus (12)
- How understanding large language models can inform the use of ChatGPT in physics education
- Unreflected Acceptance -- Investigating the Negative Consequences of ChatGPT-Assisted Problem Solving in Physics Education
- A comparison of Human, GPT-3.5, and GPT-4 Performance in a University-Level Coding Course
- Multilingual Performance of a Multimodal Artificial Intelligence System on Multisubject Physics Concept Inventories
- Assessing Confidence in AI-Assisted Grading of Physics Exams through Psychometrics: An Exploratory Study
- Ethel: A Virtual Teaching Assistant
- Performance of ChatGPT on tasks involving physics visual representations: the case of the Brief Electricity and Magnetism Assessment
- Cheat sites and artificial intelligence usage in online introductory physics courses: what is the extent and what effect does it have on assessments?
- Evaluating GPT- and Reasoning-based Large Language Models on Physics Olympiad Problems: Surpassing Human Performance and Implications for Educational Assessment
- Can ChatGPT pass a physics degree? Making a case for reformation of assessment of undergraduate degrees
- Using Large Language Models for Solving Thermodynamic Problems
- Incentivizing supplemental math assignments and using AI-generated hints is associated with improved exam performance