1 paper
Tianwei Lin, Zuyi Zhou, Xinda Zhao +6
Long-context LLM agents must access the right evidence from large environments and use it faithfully. However, the popular Needle-in-a-Haystack (NIAH) evaluation mostly measures be…