5 papers
Muse Spark Safety & Preparedness Report
Cristina Menghini, Peter Ney, Hamza Kwisaba +117
Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framewo…
Eval Factsheets: A Structured Framework for Documenting AI Evaluations
Florian Bordes, Candace Ross, Justine T Kao +2
The rapid proliferation of benchmarks has created significant challenges in reproducibility, transparency, and informed decision-making. However, unlike datasets and models -- whic…
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
Vipul Gupta, Candace Ross, David Pantoja +3
One of the most challenging problems facing NLP today is evaluation. Some of the most pressing issues pertain to benchmark saturation, data contamination, and diversity in the qual…
What makes a good metric? Evaluating automatic metrics for text-to-image consistency
Candace Ross, Melissa Hall, Adriana Romero Soriano +1
Language models are increasingly being incorporated as components in larger AI systems for various purposes, from prompt optimization to automatic evaluation. In this work, we anal…
Changing Answer Order Can Decrease MMLU Accuracy
Vipul Gupta, David Pantoja, Candace Ross +2
As large language models (LLMs) have grown in prevalence, particular benchmarks have become essential for the evaluation of these models and for understanding model capabilities. M…