When AI Evaluates Its Own Work: Validating Learner-Initiated, AI-Generated Physics Practice Problems
arXiv:2508.03085 · doi:10.1103/9hh9-vt4d
Abstract
Large language models (LLMs) can now generate physics practice problems in real time, yet the educational value of these items hinges on rapid, reliable post-generation vetting. In this exploratory study, we investigated which automated checks are both technically feasible and pedagogically meaningful when exercises are produced on demand within a chatbot interface. A cohort of 34 introductory-physics students generated and attempted 543 practice problems during exam preparation. Each item was labeled by an expert on a wide range of quality attributes and presented to the learners in pairs to record their preference. We then (i) benchmarked three commodity LLMs as ``judges'' against the expert labels, (ii) quantified which attributes predict student choice via random-forest models, and (iii) triangulated these results with free-form exit surveys. Only a small subset of the original metric items proved necessary to reliably address student preferences either directly or by proxy. The study demonstrates that scalable formative assessment does not require exhaustive scoring: a carefully curated core of structural and learner-visible checks is sufficient to ensure both technical soundness and user appeal. The findings provide a practical blueprint for deploying real-time, AI-generated practice in physics and other quantitative disciplines.
References in corpus (13)
- Survey of Hallucination in Natural Language Generation
- Could an Artificial-Intelligence agent pass an introductory physics course?
- Physics task development of prospective physics teachers using ChatGPT
- Enhancing STEM Learning with ChatGPT and Bing Chat as Objects to Think With: A Case Study
- Performance of ChatGPT on the Test of Understanding Graphs in Kinematics
- Assessing the quality of a student-generated question repository
- ChatGPT as a tool for honing teachers' Socratic dialogue skills
- Using Large Language Models to Assign Partial Credit to Students' Explanations of Problem-Solving Process: Grade at Human Level Accuracy with Grading Confidence Index and Personalized Student-facing Feedback
- Multilingual Performance of a Multimodal Artificial Intelligence System on Multisubject Physics Concept Inventories
- Assessing Confidence in AI-Assisted Grading of Physics Exams through Psychometrics: An Exploratory Study
- Ethel: A Virtual Teaching Assistant
- Can ChatGPT pass a physics degree? Making a case for reformation of assessment of undergraduate degrees
- The Boiling-Frog Problem of Physics Education