2 papers
cs.CL2026
The Trust Paradox: How CS Researchers Engage LLM Leaderboards
Pouya Sadeghi, Anamaria Crisan, Jimmy Lin
Large language model (LLM) leaderboards rank AI models using standardized benchmarks and have become highly visible across computer science, despite known limitations in their reli…
cs.CL2024
Categorical Syllogisms Revisited: A Review of the Logical Reasoning Abilities of LLMs for Analyzing Categorical Syllogism
Shi Zong, Jimmy Lin
There have been a huge number of benchmarks proposed to evaluate how large language models (LLMs) behave for logic inference tasks. However, it remains an open question how to prop…