3 papers
cs.CL2025
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
Jaime Raldua Veuthey, Zainab Ali Majid, Suhas Hariharan +1
As Large Language Models (LLMs) advance, their potential for widespread societal impact grows simultaneously. Hence, rigorous LLM evaluations are both a technical necessity and soc…
cs.AI2024
Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique
Suhas Hariharan, Zainab Ali Majid, Jaime Raldua Veuthey +1
A key development in the cybersecurity evaluations space is the work carried out by Meta, through their CyberSecEval approach. While this work is undoubtedly a useful contribution…
cs.LG2024
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
Jacob Haimes, Cenny Wenner, Kunvar Thaman +4
The training data for many Large Language Models (LLMs) is contaminated with test data. This means that public benchmarks used to assess LLMs are compromised, suggesting a performa…