1 paper
Shikhar Murty, Orr Paradise, Pratyusha Sharma
With large language models surpassing human performance on an increasing number of benchmarks, we must take a principled approach for targeted evaluation of model capabilities. Ins…