2 papers
cs.CL2025
Bayesian Evaluation of Large Language Model Behavior
Rachel Longjohn, Shang Wu, Saatvik Kher +2
It is increasingly important to evaluate how text generation systems based on large language models (LLMs) behave, such as their tendency to produce harmful output or their sensiti…
cs.LG2024
Benchmark Data Repositories for Better Benchmarking
Rachel Longjohn, Markelle Kelly, Sameer Singh +1
In machine learning research, it is common to evaluate algorithms via their performance on standard benchmark datasets. While a growing body of work establishes guidelines for -- a…