4 papers
GRADE: Generalizable Reasoning-Aware Dialogue Evaluation for AI Tutors
Parth Bhalerao, Jeromy Chang, David Chou +1
Evaluating AI tutor responses requires more than factual correctness: tutors must identify mistakes, locate errors, provide guidance, and offer actionable next steps. We present GR…
When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation
Shuowei Li, Yuming Zhao, Parth Bhalerao +1
Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored. We…
When Cultures Meet: Multicultural Text-to-Image Generation
Parth Bhalerao, Mounika Yalamarty, Brian Trinh +1
Text-to-image generation models have achieved strong performance in culturally homogeneous settings, yet their ability to generate multicultural scenes, where people and landmarks…
Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content
Parth Bhalerao, Diola Dsouza, Ruiwen Guan +1
Question answering systems are typically evaluated on factual correctness, yet many real-world applications-such as education and career guidance-require mentorship: responses that…