Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
GRADE: Generalizable Reasoning-Aware Dialogue Evaluation for AI Tutors
Parth Bhalerao, Jeromy Chang, David Chou +1
Evaluating AI tutor responses requires more than factual correctness: tutors must identify mistakes, locate errors, provide guidance, and offer actionable next steps. We present GR…
cs.CL2026
Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content
Parth Bhalerao, Diola Dsouza, Ruiwen Guan +1
Question answering systems are typically evaluated on factual correctness, yet many real-world applications-such as education and career guidance-require mentorship: responses that…