1 paper
Dongjun Kim, Gyuho Shim, Yongchan Chun +3
Large Language Models are commonly judged by their scores on standard benchmarks, yet such scores often overstate real capability since they mask the mix of skills a task actually…