1 citations · 1 across the 14 of their papers we have counts for
4 papers · 1 filter
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
Ronald Skorobogat, Ameya Prabhu, Matthias Bethge
Multilingual benchmarks guide the development of frontier models. Yet multilingual evaluations reported by frontier models are structured similar to popular reasoning and knowledge…
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
Nikhil Chandak, Shashwat Goel, Ameya Prabhu +2
Multiple choice benchmarks have long been the workhorse of language model evaluation because grading multiple choice is objective and easy to automate. However, we show multiple ch…
Data Contamination Report from the 2024 CONDA Shared Task
Oscar Sainz, Iker García-Ferrero, Alon Jacovi +25
The 1st Workshop on Data Contamination (CONDA 2024) focuses on all relevant aspects of data contamination in natural language processing, where data contamination is understood as…
CiteME: Can Language Models Accurately Cite Scientific Claims?
Ori Press, Andreas Hochlehnert, Ameya Prabhu +3
Thousands of new scientific papers are published each month. Such information overload complicates researcher efforts to stay current with the state-of-the-art as well as to verify…