4 papers
Efficient Safety Benchmarking via Item Response Theory
Fabio Spagliardi, MÃrian Silva, Ayan Datta +3
Safety benchmarks for language models are typically evaluated using static paradigms that treat all items as equally informative for all models, an assumption that is particularly…
GRAFITE: Generative Regression Analysis Framework for Issue Tracking and Evaluation
Ja Young Lee, MÃrian Silva, Mohamed Nasr +6
Large language models (LLMs) are largely motivated by their performance on popular topics and benchmarks at the time of their release. However, over time, contamination occurs due…
CARROT: A Cost Aware Rate Optimal Router
Seamus Somerstep, Felipe Maia Polo, Allysson Flavio Melo de Oliveira +5
With the rapid growth in the number of Large Language Models (LLMs), there has been a recent interest in LLM routing, or directing queries to the cheapest LLM that can deliver a su…
Efficient multi-prompt evaluation of LLMs
Felipe Maia Polo, Ronald Xu, Lucas Weber +6
Most popular benchmarks for comparing LLMs rely on a limited set of prompt templates, which may not fully capture the LLMs' abilities and can affect the reproducibility of results…