12 papers
Redact or Keep? A Fully Local AI Cascade for Educational Dialogue De-Identification
Haocheng Zhang, Zhuqian Zhou, Kirk Vanacore +2
Educational dialogue is a valuable but sensitive resource for research: the same transcripts that capture authentic learning often capture personally identifiable information (PII)…
Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset
Zhuqian Zhou, Kirk Vanacore, Bakhtawar Ahtisham +7
Large-scale sharing of dialogue data is key to advancing the science of teaching and learning, yet rigorous de-identification remains a major barrier. In mathematics tutoring trans…
Sandpiper: Orchestrated AI-Annotation for Educational Discourse at Scale
Daryl Hedley, Doug Pietrzak, Jorge Dias +9
Digital educational environments are expanding toward complex AI and human discourse, providing researchers with an abundance of data that offers deep insights into learning and in…
Million Tutoring Moves (MTM): An Open Multimodal Dataset for the Science of Tutoring
René Kizilcec, Kirk Vanacore, Zhuqian Zhou +8
We introduce the Million Tutoring Moves (MTM) project, an open dataset initiative aimed at advancing the science of tutoring through large-scale, reusable, and multimodal interacti…
Domain-Adapted Retrieval for In-Context Annotation of Pedagogical Dialogue Acts
Jinsook Lee, Kirk Vanacore, Zhuqian Zhou +2
Automated annotation of pedagogical dialogue is a high-stakes task where LLMs often fail without sufficient domain grounding. We present a domain-adapted RAG pipeline for tutoring…
Can You Tell It's AI? Human Perception of Synthetic Voices in Vishing Scenarios
Zoha Hayat Bhatti, Bakhtawar Ahtisham, Seemal Tausif +3
Large Language Models and commercial speech synthesis systems now enable highly realistic AI-generated voice scams (vishing), raising urgent concerns about deception at scale. Yet…