3 papers
cs.CL2026
The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
Li Lucy, Albert Zhang, Nathan Anderson +2
Effective mathematics education requires identifying and responding to students' mistakes. For AI to support pedagogical applications, models must perform well across different lev…
cs.LG2026
How2Everything: Mining the Web for How-To Procedures to Evaluate and Improve LLMs
Yapei Chang, Kyle Lo, Mohit Iyyer +1
Generating step-by-step "how-to" procedures is a key LLM capability: how-to advice is commonly requested in chatbots, and step-by-step planning is critical for reasoning over compl…
cs.CL2025
Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries
Chaitanya Malaviya, Joseph Chee Chang, Dan Roth +3
Language model users often issue queries that lack specification, where the context under which a query was issued -- such as the user's identity, the query's intent, and the crite…