1 paper · 1 filter
Lea Hergert, Gábor Berend, Mario Szegedy +2
Large language models (LLMs) achieve superhuman performance on complex reasoning tasks, yet often fail on much simpler problems, raising concerns about their reliability and interp…