8 papers
Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents
Wei-Chieh Huang, Weizhi Zhang, Yuchen Wu +12
Memory is becoming core infrastructure for long-horizon LLM agents, yet existing evaluations offer limited guidance on which memory substrate, namely the underlying medium in which…
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation
Yaozu Wu, Wei-Chieh Huang, Jizhou Guo +11
Large language models increasingly operate in settings where humans are active collaborators rather than passive task providers. We introduce HAS-Framework, a graph-based framework…
Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM Personalization
Linfeng Du, Ye Yuan, Zichen Zhao +8
Large language models (LLMs) excel at general-purpose tasks, yet adapting their responses to individual users remains challenging. Retrieval augmentation provides a lightweight alt…
CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification
Hanrong Zhang, Shicheng Fan, Henry Peng Zou +12
Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address. A tool is a single, self-contained func…
CARE: Privacy-Compliant Agentic Reasoning with Evidence Discordance
Haochen Liu, Weien Li, Rui Song +6
Large language model (LLM) systems are increasingly used to support high-stakes decision-making, but they typically perform worse when the available evidence is internally inconsis…
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation
Henry Peng Zou, Chunyu Miao, Wei-Chieh Huang +16
As LLM agents transition from short, static problem solving to executing complex, long-horizon tasks in dynamic environments, the ability to handle user interruptions, such as addi…