3 papers
cs.CL2026
Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset
Zhuqian Zhou, Kirk Vanacore, Bakhtawar Ahtisham +7
Large-scale sharing of dialogue data is key to advancing the science of teaching and learning, yet rigorous de-identification remains a major barrier. In mathematics tutoring trans…
cs.HC2026
Sandpiper: Orchestrated AI-Annotation for Educational Discourse at Scale
Daryl Hedley, Doug Pietrzak, Jorge Dias +9
Digital educational environments are expanding toward complex AI and human discourse, providing researchers with an abundance of data that offers deep insights into learning and in…
cs.CY2026
Million Tutoring Moves (MTM): An Open Multimodal Dataset for the Science of Tutoring
René Kizilcec, Kirk Vanacore, Zhuqian Zhou +8
We introduce the Million Tutoring Moves (MTM) project, an open dataset initiative aimed at advancing the science of tutoring through large-scale, reusable, and multimodal interacti…