1 paper · 1 filter
Ariel Procaccia, Benjamin Schiffer, Serena Wang +1
A common problem in LLM evaluation is how to choose a subset of metrics from a full suite of possible metrics. Subset selection is usually done for efficiency or interpretability r…