Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
Li Lucy, Albert Zhang, Nathan Anderson +2
Effective mathematics education requires identifying and responding to students' mistakes. For AI to support pedagogical applications, models must perform well across different lev…
cs.CL2025
Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries
Chaitanya Malaviya, Joseph Chee Chang, Dan Roth +3
Language model users often issue queries that lack specification, where the context under which a query was issued -- such as the user's identity, the query's intent, and the crite…
cs.CL2024
One Thousand and One Pairs: A "novel" challenge for long-context language models
Marzena Karpinska, Katherine Thai, Kyle Lo +2
Synthetic long-context LLM benchmarks (e.g., "needle-in-the-haystack") test only surface-level retrieval capabilities, but how well can long-context LLMs retrieve, synthesize, and…