4 papers
CREATE: Testing LLMs for Associative Creativity
Manya Wadhwa, Tiasa Singha Roy, Harvey Lederman +2
A key component of creativity is associative reasoning: the ability to draw novel yet meaningful connections between concepts. We introduce CREATE, a benchmark designed to evaluate…
Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective
Hamid Osooli, Kareema Batool, Rick Gentry +3
Weak-to-strong alignment offers a promising route to scalable supervision, but it can fail when a strong model becomes confidently wrong on examples that lie in the weak model's bl…
Interpreting and Mitigating Unwanted Uncertainty in LLMs
Tiasa Singha Roy, Ayush Rajesh Jhaveri, Ilias Triantafyllopoulos
Despite their impressive capabilities, Large Language Models (LLMs) exhibit unwanted uncertainty, a phenomenon where a model changes a previously correct answer into an incorrect o…
Can LLMs Math? -- Exploring the Pitfalls in Mathematical Reasoning
Tiasa Singha Roy, Aditeya Baral, Ayush Rajesh Jhaveri +1
Large language models (LLMs) demonstrate considerable potential in various natural language tasks but face significant challenges in mathematical reasoning, particularly in executi…