collaborators

5 papers

cs.LG2025

MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model

Prasanna Mayilvahanan, Ricardo Dominguez-Olmedo, Thaddäus Wiedemer +1

With the advent of DeepSeek-R1, a new wave of reinforcement learning (RL) methods has emerged that seem to unlock stronger mathematical reasoning. However, a closer look at the ope…

cs.LG2025

Train-before-Test Harmonizes Language Model Rankings

Guanhua Zhang, Ricardo Dominguez-Olmedo, Moritz Hardt

Existing language model benchmarks provide contradictory model rankings, even for benchmarks that aim to capture similar skills. This dilemma of conflicting rankings hampers model…

cs.CL2025

Lawma: The Power of Specialization for Legal Annotation

Ricardo Dominguez-Olmedo, Vedant Nanda, Rediet Abebe +6

Annotation and classification of legal text are central components of empirical legal research. Traditionally, these tasks are often delegated to trained research assistants. Motiv…

cs.CL2025

Training on the Test Task Confounds Evaluation and Emergence

Ricardo Dominguez-Olmedo, Florian E. Dorner, Moritz Hardt

We study a fundamental problem in the evaluation of large language models that we call training on the test task. Unlike wrongful practices like training on the test data, leakage,…

cs.CL2024

Questioning the Survey Responses of Large Language Models

Ricardo Dominguez-Olmedo, Moritz Hardt, Celestine Mendler-Dünner

Surveys have recently gained popularity as a tool to study large language models. By comparing survey responses of models to those of human reference populations, researchers aim t…