1 paper
Davis Brown, Prithvi Balehannina, Helen Jin +3
Language model evaluations often fail to characterize consequential failure modes, forcing experts to inspect outputs and build new benchmarks. We introduce task elicitation, a met…