4 papers
Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset
Zhuqian Zhou, Kirk Vanacore, Bakhtawar Ahtisham +7
Large-scale sharing of dialogue data is key to advancing the science of teaching and learning, yet rigorous de-identification remains a major barrier. In mathematics tutoring trans…
Sandpiper: Orchestrated AI-Annotation for Educational Discourse at Scale
Daryl Hedley, Doug Pietrzak, Jorge Dias +9
Digital educational environments are expanding toward complex AI and human discourse, providing researchers with an abundance of data that offers deep insights into learning and in…
Million Tutoring Moves (MTM): An Open Multimodal Dataset for the Science of Tutoring
René Kizilcec, Kirk Vanacore, Zhuqian Zhou +8
We introduce the Million Tutoring Moves (MTM) project, an open dataset initiative aimed at advancing the science of tutoring through large-scale, reusable, and multimodal interacti…
AI Annotation Orchestration: Evaluating LLM verifiers to Improve the Quality of LLM Annotations in Learning Analytics
Bakhtawar Ahtisham, Kirk Vanacore, Jinsook Lee +3
Large Language Models (LLMs) are increasingly used to annotate learning interactions, yet concerns about reliability limit their utility. We test whether verification-oriented orch…