1 paper · 1 filter
Xingyu Chen, Rui Wang, Zhaopeng Tu +1
Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specif…