13 papers
Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth
Alex Liu, Lief Esbenshade, Michael Xiao +4
Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that presumes human coding is the standard to…
Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use
Alex Liu, Min Sun, Lief Esbenshade +4
Qualitative researchers increasingly encounter interaction corpora whose scale exceeds what manual coding alone can address, and large language models (LLMs) are frequently propose…
Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering
Zewei Tian, Alex Liu, Lief Esbenshade +6
The integration of large language models (LLMs) into educational assessment represents a transformative shift in classroom grading practices. While automated scoring systems and ma…
Improving the Distributional Alignment of LLMs using Supervision
Gauri Kambhatla, Sanjana Gautam, Angela Zhang +4
The ability to accurately align LLMs with diverse population groups on subjective questions would have great value. In this work, we show that adding simple supervision can more co…
Teacher-Authored Prompts for Configuring Student-AI Dialogue: K-12 Classroom Implementation
Alex Liu, Min Sun, Lief Esbenshade +3
GenAI has rapidly entered instructional and learning settings as a teaching assistant or AI tutor. However, less is known about how pedagogical intent connects to the learning gene…
How K-12 Educators Use AI: LLM-Assisted Qualitative Analysis at Scale
Alex Liu, Lief Esbenshade, Shawon Sarkar +4
This study investigates how K-12 educators use generative AI tools in real-world instructional contexts and how large language models (LLMs) can support scalable qualitative analys…