1 paper · 1 filter
Ken Gu, Advait Bhat, Mike A Merrill +4
Evaluating the reasoning ability of language models (LMs) is complicated by their extensive parametric world knowledge, where benchmark performance often reflects factual recall ra…