activity
20242026
collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry

Hsien Xin Peng, Anthony Kim, Alvin Li +3

Foundation models such as GPT and Claude now solve olympiad-level mathematics with remarkable proficiency, so much so that geometry problem solving has become a standard proxy for…

cs.AI2026

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

Yash Tripathi, Silu Sharma, Sai Sidhanth Manoharan Jayanthi +2

Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce IntegrityBe…

cs.AI2026

Poker Arena: Multi-Axis Profiling of Strategic Reasoning and Memory in LLMs

Pratham Singla, Shivank Garg, Vihan Singh

Strategic reasoning under uncertainty underpins consequential decisions in negotiation, finance, and policy, but prevailing game-play benchmarks collapse heterogeneous reasoning di…

cs.AI2026

Recurrent Reasoning on Symbolic Puzzles with Sequence Models

Gowrav Mannem, Chowdhury Marzia Mahjabin, Jason Chen +2

Large language models often appear strong on symbolic and algorithmic tasks, yet this apparent strength can hide brittle behaviour when problems become longer, harder, or slightly…

cs.AI2026

SIDiffAgent: Self-Improving Diffusion Agent

Shivank Garg, Ayush Singh, Gaurav Kumar Nayak

Text-to-image diffusion models have revolutionized generative AI, enabling high-quality and photorealistic image synthesis. However, their practical deployment remains hindered by…