1 paper · 1 filter
Xingyu Chen, Rui Wang, Zhaopeng Tu +1
Fixed benchmarks are costly to renew and cannot adapt their questions to model-specific failures. We ask whether LLMs can instead discover one another's weaknesses and turn those o…