How understanding large language models can inform the use of ChatGPT in physics education
arXiv:2309.12074 · doi:10.1088/1361-6404/ad1420
Abstract
The paper aims to fulfil three main functions: (1) to serve as an introduction for the physics education community to the functioning of Large Language Models (LLMs), (2) to present a series of illustrative examples demonstrating how prompt-engineering techniques can impact LLMs performance on conceptual physics tasks and (3) to discuss potential implications of the understanding of LLMs and prompt engineering for physics teaching and learning. We first summarise existing research on the performance of a popular LLM-based chatbot (ChatGPT) on physics tasks. We then give a basic account of how LLMs work, illustrate essential features of their functioning, and discuss their strengths and limitations. Equipped with this knowledge, we discuss some challenges with generating useful output with ChatGPT-4 in the context of introductory physics, paying special attention to conceptual questions and problems. We then provide a condensed overview of relevant literature on prompt engineering and demonstrate through illustrative examples how selected prompt-engineering techniques can be employed to improve ChatGPT-4's output on conceptual introductory physics problems. Qualitatively studying these examples provides additional insights into ChatGPT's functioning and its utility in physics problem solving. Finally, we consider how insights from the paper can inform the use of LLMs in the teaching and learning of physics.
References in corpus (32)
- Survey of Hallucination in Natural Language Generation
- Training language models to follow instructions with human feedback
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- A Survey of Large Language Models
- Large Language Models are Zero-Shot Reasoners
- Emergent Abilities of Large Language Models
- A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Capabilities of GPT-4 on Medical Challenge Problems
- Gender bias and stereotypes in Large Language Models
- Mathematical Capabilities of ChatGPT
- Measuring Mathematical Problem Solving With the MATH Dataset
- The Death of the Short-Form Physics Essay in the Coming AI Revolution
- Could an Artificial-Intelligence agent pass an introductory physics course?
- Enhancing STEM Learning with ChatGPT and Bing Chat as Objects to Think With: A Case Study
- Large Language Models Can Be Easily Distracted by Irrelevant Context
- How do physics students evaluate artificial intelligence responses on comprehension questions? A study on the perceived scientific accuracy and linguistic quality
- Eight Things to Know about Large Language Models
- How Context Affects Language Models' Factual Predictions
- The Impact of AI in Physics Education: A Comprehensive Review from GCSE to University Levels
- In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPT
- Unreflected Acceptance -- Investigating the Negative Consequences of ChatGPT-Assisted Problem Solving in Physics Education
- ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
- Aligning Large Language Models with Human: A Survey
- Large Language Model Guided Tree-of-Thought
- PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change
- AI and the FCI: Can ChatGPT Project an Understanding of Introductory Physics?
- Advances in apparent conceptual physics reasoning in GPT-4
- Unveiling Gender Bias in Terms of Profession Across LLMs: Analyzing and Addressing Sociological Implications
- Exploring Durham University Physics exams with Large Language Models
Cited by in corpus (13)
- Exploring Generative AI assisted feedback writing for students' written responses to a physics conceptual question with prompt engineering and few-shot learning
- Performance of ChatGPT on the Test of Understanding Graphs in Kinematics
- A comparison of Human, GPT-3.5, and GPT-4 Performance in a University-Level Coding Course
- ChatGPT as a tool for honing teachers' Socratic dialogue skills
- Multilingual Performance of a Multimodal Artificial Intelligence System on Multisubject Physics Concept Inventories
- Assessing Confidence in AI-Assisted Grading of Physics Exams through Psychometrics: An Exploratory Study
- Evaluating vision-capable chatbots in interpreting kinematics graphs: a comparative study of free and subscription-based models
- Performance of ChatGPT on tasks involving physics visual representations: the case of the Brief Electricity and Magnetism Assessment
- Evaluating GPT- and Reasoning-based Large Language Models on Physics Olympiad Problems: Surpassing Human Performance and Implications for Educational Assessment
- Can ChatGPT pass a physics degree? Making a case for reformation of assessment of undergraduate degrees
- Using AI Large Language Models for Grading in Education: A Hands-On Test for Physics
- Creating a customisable Socratic AI physics tutor
- Exploring Large Language Models (LLMs) through interactive Python activities