2 papers
cs.CL2024
REBUS: A Robust Evaluation Benchmark of Understanding Symbols
Andrew Gritsevskiy, Arjun Panickssery, Aaron Kirtland +7
We propose a new benchmark evaluating the performance of multimodal large language models on rebus puzzles. The dataset covers 333 original examples of image-based wordplay, cluing…
cs.CL2024
Inverse Scaling: When Bigger Isn't Better
Ian R. McKenzie, Alexander Lyzhov, Michael Pieler +24
Work on scaling laws has found that large language models (LMs) show predictable improvements to overall loss with increased scale (model size, training data, and compute). Here, w…