Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
When Verification Hurts: Asymmetric Effects of Multi-Agent Feedback in Logic Proof Tutoring
Tahreem Yasir, Sutapa Dey Tithi, Benyamin Tabarsi +7
Large language models (LLMs) are increasingly used for automated tutoring, but their reliability in structured symbolic domains remains unclear. We study step-level feedback for pr…
cs.AI2026
Enhancing Mathematical Problem Solving in LLMs through Execution-Driven Reasoning Augmentation
Aditya Basarkar, Benyamin Tabarsi, Tiffany Barnes +1
Mathematical problem solving is a fundamental benchmark for assessing the reasoning capabilities of artificial intelligence and a gateway to applications in education, science, and…