1 citations · 1 across the 3 of their papers we have counts for
1 paper · 2 filters
Natasha Butt, Blazej Manczak, Auke Wiggers +4
Large language models are increasingly solving tasks that are commonly believed to require human-level reasoning ability. However, these models still perform very poorly on benchma…