From the 1 of 11 linked papers with an AI index.
2 citations · 2 across the 5 of their papers we have counts for
11 papers
Pigeonholing: how bad prompts hurt models, causing collapse and mistakes
Hyunji Nam, Keertana Chidambaram, Dorottya Demszky +1
The paper studies how poorly chosen prompts and conversation contexts cause large language models to repeat mistakes, narrow their output diversity, and change stances—a problem th…
EduCoder: An Open-Source Annotation System for Education Transcript Data
Saad Ashraf, James Malamut, Vishal Kumar +7
We introduce EduCoder, a domain-specialized tool designed to support utterance-level annotation of educational dialogue. While general-purpose text annotation tools for NLP and qua…
From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment
Ivo Bueno, Babette Bühler, Philipp Stark +5
Automated scoring models are increasingly used to assign rubric-based quality ratings to complex language performances, including classroom transcripts, yet they typically provide…
Mitigating LLM biases toward spurious social contexts using direct preference optimization
Hyunji Nam, Dorottya Demszky
LLMs are increasingly used for high-stakes decision-making, yet their sensitivity to spurious contextual information can introduce harmful biases. This is a critical concern when m…
Practitioner Voices Summit: How Teachers Evaluate AI Tools through Deliberative Sensemaking
Dorottya Demszky, Christopher Mah, Helen Higgins
Teachers face growing pressure to integrate AI tools into their classrooms, yet are rarely positioned as agentic decision-makers in this process. Understanding the criteria teacher…
Marked Pedagogies: Examining Linguistic Biases in Personalized Automated Writing Feedback
Mei Tan, Lena Phalen, Dorottya Demszky
Effective personalized feedback is critical to students' literacy development. Though LLM-powered tools now promise to automate such feedback at scale, LLMs are not language-neutra…