2 papers
cs.AI2025
Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language Models
Simeng Han, Howard Dai, Stephen Xia +7
Accuracy remains a standard metric for evaluating AI systems, but it offers limited insight into how models arrive at their solutions. In this work, we introduce a benchmark based…
cs.CL2025
Learning to Reason via Mixture-of-Thought for Logical Reasoning
Tong Zheng, Lichang Chen, Simeng Han +2
Human beings naturally utilize multiple reasoning modalities to learn and solve logical problems, i.e., different representational formats such as natural language, code, and symbo…