3 papers
cs.AI2026
When Verification Hurts: Asymmetric Effects of Multi-Agent Feedback in Logic Proof Tutoring
Tahreem Yasir, Sutapa Dey Tithi, Benyamin Tabarsi +7
Large language models (LLMs) are increasingly used for automated tutoring, but their reliability in structured symbolic domains remains unclear. We study step-level feedback for pr…
cs.AI2026
Enhancing Mathematical Problem Solving in LLMs through Execution-Driven Reasoning Augmentation
Aditya Basarkar, Benyamin Tabarsi, Tiffany Barnes +1
Mathematical problem solving is a fundamental benchmark for assessing the reasoning capabilities of artificial intelligence and a gateway to applications in education, science, and…
cs.CL2026
SafeTalkCoach: Diversity-Driven Multi-Agent Simulation for Parent-Teen Health Conversations
Benyamin Tabarsi, Wenbo Li, Tahreem Yasir +4
The importance of effective parent-child communication about sexual health is widely acknowledged, but real-world data on these conversations is scarce and challenging to collect,…