Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
Hui Dai, Ryan Teehan, Mengye Ren
Many existing evaluation benchmarks for Large Language Models (LLMs) quickly become outdated due to the emergence of new models and training data. These benchmarks also fall short…
cs.CL2024
DENIAHL: In-Context Features Influence LLM Needle-In-A-Haystack Abilities
Hui Dai, Dan Pechi, Xinyi Yang +2
The Needle-in-a-haystack (NIAH) test is a general task used to assess language models' (LMs') abilities to recall particular information from long input context. This framework how…