1 paper
Howard Yen, Tianyu Gao, Minmin Hou +5
Many benchmarks exist for evaluating long-context language models (LCLMs), yet developers often rely on synthetic tasks such as needle-in-a-haystack (NIAH) or an arbitrary subset o…