Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
Jon Saad-Falcon, Rajan Vivek, William Berrios +6
As language models become integral to critical workflows, assessing their behavior remains a fundamental challenge -- human evaluation is costly and noisy, while automated metrics…
cs.CL2024
Anchor Points: Benchmarking Models with Much Fewer Examples
Rajan Vivek, Kawin Ethayarajh, Diyi Yang +1
Modern language models often exhibit powerful but brittle behavior, leading to the development of larger and more diverse benchmarks to reliably assess their behavior. Here, we sug…