activity
20242026
collaborators

5 papers

cs.CY2026

Muse Spark Safety & Preparedness Report

Cristina Menghini, Peter Ney, Hamza Kwisaba +117

Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framewo…

cs.LG2025

Eval Factsheets: A Structured Framework for Documenting AI Evaluations

Florian Bordes, Candace Ross, Justine T Kao +2

The rapid proliferation of benchmarks has created significant challenges in reproducibility, transparency, and informed decision-making. However, unlike datasets and models -- whic…

cs.CL2025

Improving Model Evaluation using SMART Filtering of Benchmark Datasets

Vipul Gupta, Candace Ross, David Pantoja +3

One of the most challenging problems facing NLP today is evaluation. Some of the most pressing issues pertain to benchmark saturation, data contamination, and diversity in the qual…

cs.CL2024

What makes a good metric? Evaluating automatic metrics for text-to-image consistency

Candace Ross, Melissa Hall, Adriana Romero Soriano +1

Language models are increasingly being incorporated as components in larger AI systems for various purposes, from prompt optimization to automatic evaluation. In this work, we anal…

cs.CL2024

Changing Answer Order Can Decrease MMLU Accuracy

Vipul Gupta, David Pantoja, Candace Ross +2

As large language models (LLMs) have grown in prevalence, particular benchmarks have become essential for the evaluation of these models and for understanding model capabilities. M…