1 paper
Qiaoyuan Zheng, Yiqu Yang
Small leaderboard gaps are often interpreted as evidence that one language model is better than another, but their sign may depend on which benchmark items are included. We test th…