1 paper · 1 filter
Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza
Large language models now score near ceiling on general benchmarks, but these aggregate measures reveal little about how models behave within single disciplines. Existing art-focus…