1 paper
Tadanobu Chuyo Kamijo, Ori Rottenstreich, Javier Conde +2
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimize…