A Large Language Model Approach to Educational Survey Feedback Analysis
arXiv:2309.17447 · doi:10.1007/s40593-024-00414-0
Abstract
This paper assesses the potential for the large language models (LLMs) GPT-4 and GPT-3.5 to aid in deriving insight from education feedback surveys. Exploration of LLM use cases in education has focused on teaching and learning, with less exploration of capabilities in education feedback analysis. Survey analysis in education involves goals such as finding gaps in curricula or evaluating teachers, often requiring time-consuming manual processing of textual responses. LLMs have the potential to provide a flexible means of achieving these goals without specialized machine learning models or fine-tuning. We demonstrate a versatile approach to such goals by treating them as sequences of natural language processing (NLP) tasks including classification (multi-label, multi-class, and binary), extraction, thematic analysis, and sentiment analysis, each performed by LLM. We apply these workflows to a real-world dataset of 2500 end-of-course survey comments from biomedical science courses, and evaluate a zero-shot approach (i.e., requiring no examples or labeled training data) across all tasks, reflecting education settings, where labeled data is often scarce. By applying effective prompting practices, we achieve human-level performance on multiple tasks with GPT-4, enabling workflows necessary to achieve typical goals. We also show the potential of inspecting LLMs' chain-of-thought (CoT) reasoning for providing insight that may foster confidence in practice. Moreover, this study features development of a versatile set of classification categories, suitable for various course types (online, hybrid, or in-person) and amenable to customization. Our results suggest that LLMs can be used to derive a range of insights from survey text.
References in corpus (16)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
- A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- ReAct: Synergizing Reasoning and Acting in Language Models
- A Review of the Trends and Challenges in Adopting Natural Language Processing Methods for Education Feedback Analysis
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
- Self-Refine: Iterative Refinement with Self-Feedback
- Is ChatGPT better than Human Annotators? Potential and Limitations of ChatGPT in Explaining Implicit Hate Speech
- Sentiment analysis and opinion mining on educational data: A survey
- Instruction Tuning with GPT-4
- ChatGPT-4 Outperforms Experts and Crowd Workers in Annotating Political Twitter Messages with Zero-Shot Learning
- Efficient Few-Shot Learning Without Prompts
- Artificial Artificial Artificial Intelligence: Crowd Workers Widely Use Large Language Models for Text Production Tasks
- TimeLMs: Diachronic Language Models from Twitter
- Extracting Self-Consistent Causal Insights from Users Feedback with LLMs and In-context Learning