1 paper · 1 filter
Gonçalo Hora de Carvalho, Oscar Knap, Robert Pollice
We developed a benchmark set to assess the generalization of state-of-the-art large language models on problems beyond linguistic tasks and evaluate it on a systematic progression…