1 citations · 1 across the 13 of their papers we have counts for
4 papers · 1 filter
Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry
Hsien Xin Peng, Anthony Kim, Alvin Li +3
Foundation models such as GPT and Claude now solve olympiad-level mathematics with remarkable proficiency, so much so that geometry problem solving has become a standard proxy for…
Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists
Yash Tripathi, Silu Sharma, Sai Sidhanth Manoharan Jayanthi +2
Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce IntegrityBe…
Poker Arena: Multi-Axis Profiling of Strategic Reasoning and Memory in LLMs
Pratham Singla, Shivank Garg, Vihan Singh
Strategic reasoning under uncertainty underpins consequential decisions in negotiation, finance, and policy, but prevailing game-play benchmarks collapse heterogeneous reasoning di…
SIDiffAgent: Self-Improving Diffusion Agent
Shivank Garg, Ayush Singh, Gaurav Kumar Nayak
Text-to-image diffusion models have revolutionized generative AI, enabling high-quality and photorealistic image synthesis. However, their practical deployment remains hindered by…