12 papers · 1 filter
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
Heejin Do, Shashank Sonkar, Mrinmaya Sachan
Large language models (LLMs) can fluently generate student-like responses, making them attractive as simulated students for training and evaluating AI tutors and human educators. Y…
When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
Robinson Ferrer, Damla Turgut, Zhongzhou Chen +1
Large Language Models (LLMs) show promise for automated grading, but their outputs can be unreliable. Rather than improving grading accuracy directly, we address a complementary pr…
Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation
Yanick Zengaffinen, Andreas Opedal, Donya Rooein +3
Modeling plausible student misconceptions is critical for AI in education. In this work, we examine how large language models (LLMs) reason about misconceptions when generating mul…
MalruleLib: Large-Scale Executable Misconception Reasoning with Step Traces for Modeling Student Thinking in Mathematics
Xinghe Chen, Naiming Liu, Shashank Sonkar
Student mistakes in mathematics are often systematic: a learner applies a coherent but wrong procedure and repeats it across contexts. We introduce MalruleLib, a learning-science-g…
CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models
Naiming Liu, Richard Baraniuk, Shashank Sonkar
We introduce CLEAR-3K, a dataset of 3,000 assertion-reasoning questions designed to evaluate whether language models can determine if one statement causally explains another. Each…
Do LLMs Make Mistakes Like Students? Exploring Natural Alignment between Language Models and Human Error Patterns
Naiming Liu, Shashank Sonkar, Richard G. Baraniuk
Large Language Models (LLMs) have demonstrated remarkable capabilities in various educational tasks, yet their alignment with human learning patterns, particularly in predicting wh…