6 papers · 1 filter
Beyond Accuracy: Diagnosing Algebraic Reasoning Failures in LLMs Across Nine Complexity Dimensions
Parth Patil, Dhruv Kumar, Yash Sinha +1
Algebraic reasoning remains one of the most informative stress tests for large language models, yet current benchmarks provide no mechanism for attributing failure to a specific ca…
Measuring Representation Robustness in Large Language Models for Geometry
Vedant Jawandhia, Yash Sinha, Murari Mandal +2
Large language models (LLMs) are increasingly evaluated on mathematical reasoning, yet their robustness to equivalent problem representations remains poorly understood. In geometry…
The Compliance Paradox: Semantic-Instruction Decoupling in Automated Academic Code Evaluation
Devanshu Sahoo, Manish Prasad, Vasudev Majhi +5
The rapid integration of Large Language Models (LLMs) into educational assessment rests on the unverified assumption that instruction following capability translates directly to ob…
Confidence is Not Competence
Debdeep Sanyal, Manya Pandey, Dhruv Kumar +2
Large language models (LLMs) often exhibit a puzzling disconnect between their asserted confidence and actual problem-solving competence. We offer a mechanistic account of this dec…
Policy Optimization Prefers The Path of Least Resistance
Debdeep Sanyal, Aakash Sen Sharma, Dhruv Kumar +2
Policy optimization (PO) algorithms are used to refine Large Language Models for complex, multi-step reasoning. Current state-of-the-art pipelines enforce a strict think-then-answe…
ReviewEval: An Evaluation Framework for AI-Generated Reviews
Madhav Krishan Garg, Tejash Prasad, Tanmay Singhal +3
The escalating volume of academic research, coupled with a shortage of qualified reviewers, necessitates innovative approaches to peer review. In this work, we propose: 1. ReviewEv…