9 papers
Redact or Keep? A Fully Local AI Cascade for Educational Dialogue De-Identification
Haocheng Zhang, Zhuqian Zhou, Kirk Vanacore +2
Educational dialogue is a valuable but sensitive resource for research: the same transcripts that capture authentic learning often capture personally identifiable information (PII)…
Sandpiper: Orchestrated AI-Annotation for Educational Discourse at Scale
Daryl Hedley, Doug Pietrzak, Jorge Dias +9
Digital educational environments are expanding toward complex AI and human discourse, providing researchers with an abundance of data that offers deep insights into learning and in…
Million Tutoring Moves (MTM): An Open Multimodal Dataset for the Science of Tutoring
René Kizilcec, Kirk Vanacore, Zhuqian Zhou +8
We introduce the Million Tutoring Moves (MTM) project, an open dataset initiative aimed at advancing the science of tutoring through large-scale, reusable, and multimodal interacti…
Domain-Adapted Retrieval for In-Context Annotation of Pedagogical Dialogue Acts
Jinsook Lee, Kirk Vanacore, Zhuqian Zhou +2
Automated annotation of pedagogical dialogue is a high-stakes task where LLMs often fail without sufficient domain grounding. We present a domain-adapted RAG pipeline for tutoring…
Tutor Move Taxonomy: A Theory-Aligned Framework for Analyzing Instructional Moves in Tutoring
Zhuqian Zhou, Kirk Vanacore, Tamisha Thompson +2
Understanding what makes tutoring effective requires methods for systematically analyzing tutors' instructional actions during learning interactions. This paper presents a tutor mo…
LLM Reasoning Predicts When Models Are Right: Evidence from Coding Classroom Discourse
Bakhtawar Ahtisham, Kirk Vanacore, Zhuqian Zhou +2
Large Language Models (LLMs) are increasingly deployed to automatically label and analyze educational dialogue at scale, yet current pipelines lack reliable ways to detect when mod…