4 citations · 4 across the 2 of their papers we have counts for
4 papers
Beyond the Answer Key: Robustness Evaluation of Large Language Models for Step-Level Mathematical Verification
Fateme Mazdarani, Carlos Toxtli
Large language models (LLMs) are increasingly used as graders, verifiers, and process auditors, but most mathematical evaluations still emphasize final-answer accuracy. This can ob…
TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations
Fateme Mazdarani, Carlos Toxtli
AI systems increasingly operate between flexible input representations and formal objects used by downstream tools. A key challenge is recognizing when an unfamiliar formulation de…
Human-Centered Automation
Carlos Toxtli
The rapid advancement of Generative Artificial Intelligence (AI), such as Large Language Models (LLMs) and Multimodal Large Language Models (MLLM), has the potential to revolutioni…
Conceptual Framework for Autonomous Cognitive Entities
David Shapiro, Wangfan Li, Manuel Delaflor +1
The rapid development and adoption of Generative AI (GAI) technology in the form of chatbots such as ChatGPT and Claude has greatly increased interest in agentic machines. This pap…