25 papers
DynamicMem: A Long-Horizon Memory Benchmark in Real-World Settings
Wenya Xie, Shengming Zhou, Zelin Li +9
LLM agents increasingly act as personal assistants that must remember a user's profile over months: who they are (attributes), what they routinely do (habits), and what they prefer…
LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories
Zirui He, Haiyan Zhao, Yingcong Li +2
Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, where benchmark contamination can make memorized solutions appear as genuine…
Rep2Text: Decoding Full Text from a Single LLM Token Representation
Haiyan Zhao, Zirui He, Yiming Tang +4
Large language models (LLMs) have achieved remarkable progress across diverse tasks, yet their internal mechanisms remain largely opaque. In this work, we investigate a fundamental…
FAMA: Failure-Aware Meta-Agentic Framework for Open-Source LLMs in Interactive Tool Use Environments
Amir Saeidi, Venkatesh Mishra, Souradeep Mukhopadhyay +4
Large Language Models are being increasingly deployed as the decision-making core of autonomous agents capable of effecting change in external environments. Yet, in conversational…
Scaling Teams or Scaling Time? Memory Enabled Lifelong Learning in LLM Multi-Agent Systems
Shanglin Wu, Yuyang Luo, Yueqing Liang +4
Large language model (LLM) multi-agent systems can scale along two distinct dimensions: by increasing the number of agents and by improving through accumulated experience over time…
A Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon Agents
Wei-Chieh Huang, Weizhi Zhang, Yueqing Liang +57
Research in artificial intelligence is shifting from model innovations and benchmark scores towards problem definition and rigorous real-world evaluation. As the field enters the "…