Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
BIG-Bench Extra Hard
Mehran Kazemi, Bahare Fatemi, Hritik Bansal +17
Large language models (LLMs) are increasingly deployed in everyday applications, demanding robust general reasoning capabilities and diverse reasoning skillset. However, current LL…
cs.CL2024
Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models
Piotr Padlewski, Max Bain, Matthew Henderson +19
We introduce Vibe-Eval: a new open benchmark and framework for evaluating multimodal chat models. Vibe-Eval consists of 269 visual understanding prompts, including 100 of hard diff…