1 paper · 1 filter
Yixiong Fang, Tianran Sun, Yuling Shi +2
The increasing complexity of large language models (LLMs) raises concerns about their ability to "cheat" on standard Question Answering (QA) benchmarks by memorizing task-specific…