5 citations · 5 across the 1 of their papers we have counts for
2 papers
cs.CL2025
Bayesian Evaluation of Large Language Model Behavior
Rachel Longjohn, Shang Wu, Saatvik Kher +2
It is increasingly important to evaluate how text generation systems based on large language models (LLMs) behave, such as their tendency to produce harmful output or their sensiti…
cs.LG2024★ 5 cited
Benchmark Data Repositories for Better Benchmarking
Rachel Longjohn, Markelle Kelly, Sameer Singh +1
In machine learning research, it is common to evaluate algorithms via their performance on standard benchmark datasets. While a growing body of work establishes guidelines for -- a…