8 papers
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Mubashara Akhtar, Anka Reuel, Prajna Soni +36
Artificial intelligence benchmarks are an important mechanism to measure model progress and guide deployment decisions. However, benchmarks quickly "saturate", making it difficult…
Not What, But How: A Framework for Auditing LLM Responses across Positioning, Generalization, Anthropomorphism, and Maxims
Siddhesh Milind Pawar, Sarah Masud, Haneul Yoo +2
Large language models (LLMs) are being increasingly used to answer subjective, information-seeking questions, where users are sensitive to how responses are communicated, not just…
Capabilities Ain't All You Need: Measuring Propensities in AI
Daniel Romero-Alvarado, Fernando MartÃnez-Plumed, Lorenzo Pacchiardi +11
AI evaluation has primarily focused on measuring capabilities, with formal approaches inspired from Item Response Theory (IRT) being increasingly applied. Yet propensities - the te…
BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar +3
Understanding biases and stereotypes encoded in the weights of Large Language Models (LLMs) is crucial for developing effective mitigation strategies. However, biased behaviour is…
CulTrace: Tracing Internal Cultural Reasoning in Large Language Models
Haeun Yu, Arnav Arora Seogyeong Jeong, Seogyeong Jeong +8
The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of models' hidden representations of different cultures.…
Evaluation of Cultural Competence of Vision-Language Models
Srishti Yadav, Lauren Tilton, Maria Antoniak +10
Modern vision-language models (VLMs) often fail at cultural competency evaluations and benchmarks. Given the diversity of applications built upon VLMs, there is renewed interest in…