1 paper
Yangzhen Wu, Aaron J. Li, Wenjie Ma +10
The rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the ability of existing datasets to differentiate model capabilities or pr…