15 papers · 1 filter
DynamicMem: A Long-Horizon Memory Benchmark in Real-World Settings
Wenya Xie, Shengming Zhou, Zelin Li +9
LLM agents increasingly act as personal assistants that must remember a user's profile over months: who they are (attributes), what they routinely do (habits), and what they prefer…
LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories
Zirui He, Haiyan Zhao, Yingcong Li +2
Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, where benchmark contamination can make memorized solutions appear as genuine…
Rep2Text: Decoding Full Text from a Single LLM Token Representation
Haiyan Zhao, Zirui He, Yiming Tang +4
Large language models (LLMs) have achieved remarkable progress across diverse tasks, yet their internal mechanisms remain largely opaque. In this work, we investigate a fundamental…
FAMA: Failure-Aware Meta-Agentic Framework for Open-Source LLMs in Interactive Tool Use Environments
Amir Saeidi, Venkatesh Mishra, Souradeep Mukhopadhyay +4
Large Language Models are being increasingly deployed as the decision-making core of autonomous agents capable of effecting change in external environments. Yet, in conversational…
A Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon Agents
Wei-Chieh Huang, Weizhi Zhang, Yueqing Liang +57
Research in artificial intelligence is shifting from model innovations and benchmark scores towards problem definition and rigorous real-world evaluation. As the field enters the "…
Benchmarking LLMs for Political Science: A United Nations Perspective
Yueqing Liang, Liangwei Yang, Chen Wang +6
Large Language Models (LLMs) have achieved significant advances in natural language processing, yet their potential for high-stake political decision-making remains largely unexplo…