14 papers
Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry
Hsien Xin Peng, Anthony Kim, Alvin Li +3
Foundation models such as GPT and Claude now solve olympiad-level mathematics with remarkable proficiency, so much so that geometry problem solving has become a standard proxy for…
Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists
Yash Tripathi, Silu Sharma, Sai Sidhanth Manoharan Jayanthi +2
Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce IntegrityBe…
Poker Arena: Multi-Axis Profiling of Strategic Reasoning and Memory in LLMs
Pratham Singla, Shivank Garg, Vihan Singh
Strategic reasoning under uncertainty underpins consequential decisions in negotiation, finance, and policy, but prevailing game-play benchmarks collapse heterogeneous reasoning di…
Do Vision-Language Models See or Guess? Measuring and Reducing Textual-Prior Reliance with a Phrasing-Controlled Benchmark
Pratham Singla, Shivank Garg, Vihan Singh +1
Vision-language models (VLMs) are increasingly deployed where answers must follow from what is in the image, yet they often answer from textual priors, the question's phrasing toge…
Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
Pratham Singla, Shivank Garg, Ayush Singh +2
Recent advances in post-training techniques have endowed Large Language Models (LLMs) with enhanced capabilities for tackling complex, logic-intensive tasks through the generation…
Recurrent Reasoning on Symbolic Puzzles with Sequence Models
Gowrav Mannem, Chowdhury Marzia Mahjabin, Jason Chen +2
Large language models often appear strong on symbolic and algorithmic tasks, yet this apparent strength can hide brittle behaviour when problems become longer, harder, or slightly…