1 paper
Kyle Richardson, Cullen Anderson, Pranav Balakrishnan +4
While Large Language Models have improved rapidly, many fundamental questions remain about how to evaluate the knowledge and reasoning abilities they acquire, and how such evaluati…