17 citations · 33 across the 6 of their papers we have counts for
1 paper · 1 filter
Natasha Butt, Blazej Manczak, Auke Wiggers +4
Large language models are increasingly solving tasks that are commonly believed to require human-level reasoning ability. However, these models still perform very poorly on benchma…