3 papers
cs.CL2026
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?
Hanxiang Chao, Yihan Bai, Rui Sheng +2
Large Language Model (LLM) agents are increasingly expected to maintain coherent, long-term personalized memory, yet current benchmarks primarily measure static fact retrieval, ove…
cs.CL2026
Search Arena: Analyzing Search-Augmented LLMs
Mihran Miroyan, Tsung-Han Wu, Logan King +8
Search-augmented language models combine web search with Large Language Models (LLMs) to improve response groundedness and freshness. However, analyzing these systems remains chall…
cs.LG2025
Prompt-to-Leaderboard
Evan Frick, Connor Chen, Joseph Tennyson +4
Large language model (LLM) evaluations typically rely on aggregated metrics like accuracy or human preference, averaging across users and prompts. This averaging obscures user- and…