1 paper
Jon Saad-Falcon, Rajan Vivek, William Berrios +6
As language models become integral to critical workflows, assessing their behavior remains a fundamental challenge -- human evaluation is costly and noisy, while automated metrics…