4 citations · 4 across the 2 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Beyond the Answer Key: Robustness Evaluation of Large Language Models for Step-Level Mathematical Verification
Fateme Mazdarani, Carlos Toxtli
Large language models (LLMs) are increasingly used as graders, verifiers, and process auditors, but most mathematical evaluations still emphasize final-answer accuracy. This can ob…
cs.AI2026
TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations
Fateme Mazdarani, Carlos Toxtli
AI systems increasingly operate between flexible input representations and formal objects used by downstream tools. A key challenge is recognizing when an unfamiliar formulation de…