2 papers
physics.ed-ph2026
Feedback Indices to Evaluate LLM Responses to Rebuttals for Multiple Choice Type Questions
Justin C. Dunlap, Anne-Simone Parent, Ralf Widenhorn
We present a systematic framework of indices designed to characterize Large Language Model (LLM) responses when challenged with rebuttals during a chat. Assessing how LLMs respond…
physics.ed-ph2025
Translating the Force Concept Inventory in the age of AI
Marina Babayeva, Justin Dunlap, Marie Snětinová +1
We present a study that translates the Force Concept Inventory (FCI) using OpenAI GPT-4o and assess the specific difficulties of translating a scientific-focused topic using Large…