4 papers
Reason-Mediated Behavioral Models for Auditing LLM Social Simulators
Atharva Pandey, Gautam Jajoo
Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether simulated outcomes resemble human outcomes…
On the Internal Semantics of Time-Series Foundation Models
Atharva Pandey, Abhilash Neog, Gautam Jajoo
Time-series Foundation Models (TSFMs) have recently emerged as a universal paradigm for learning across diverse temporal domains. However, despite their empirical success, the inte…
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
Xinnuo Xu, Rachel Lawrence, Kshitij Dubey +7
Recent Large Language Models (LLMs) have reported high accuracy on reasoning benchmarks. However, it is still unclear whether the observed results arise from true reasoning or from…
DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
Atharva Pandey, Kshitij Dubey, Rahul Sharma +1
Despite great performance on Olympiad-level reasoning problems, frontier large language models can still struggle on high school math when presented with novel problems outside sta…