1 citations · 1 across the 3 of their papers we have counts for
3 papers
HARP: A challenging human-annotated math reasoning benchmark
Albert S. Yue, Lovish Madaan, Ted Moskovitz +2
Math reasoning is becoming an ever increasing area of focus as we scale large language models. However, even the previously-toughest evals like MATH are now close to saturated by f…
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
Lovish Madaan, David Esiobu, Pontus Stenetorp +2
In the recent past, a popular way of evaluating natural language understanding (NLU), was to consider a model's ability to perform natural language inference (NLI) tasks. In this p…
Quantifying Variance in Evaluation Benchmarks
Lovish Madaan, Aaditya K. Singh, Rylan Schaeffer +5
Evaluation benchmarks are the cornerstone of measuring capabilities of large language models (LLMs), as well as driving progress in said capabilities. Originally designed to make c…