4 papers
Measuring Intelligence Beyond Human Scale
Jerry Han, Rafael Moschopoulos, Ella Colby +5
How can we measure intelligence beyond human capability? Human-authored benchmarks saturate, and above human capability, examiners may not know which tasks are both hard and verifi…
GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models
Vartan Shadarevian, Kia Ghods, Alex Kenich +1
Large language models (LLMs) are increasingly deployed as economic agents in marketplaces, auctions, and bidding settings. Anticipating their behavior in any specific deployment is…
Visual serial processing deficits explain divergences in human and VLM reasoning
Nicholas Budny, Kia Ghods, Declan Campbell +6
Why do Vision Language Models (VLMs), despite success on standard benchmarks, often fail to match human performance on surprisingly simple visual reasoning tasks? While the underly…
Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem
Declan Campbell, Sunayana Rane, Tyler Giallanza +8
Recent work has documented striking heterogeneity in the performance of state-of-the-art vision language models (VLMs), including both multimodal language models and text-to-image…