1 paper
Alex Pan, Mary-Anne Williams
The dominant way of judging Large Language Models (LLMs) has been to ask how well they can recall explicit facts from very long inputs. While today's best models achieve near perfe…