5 papers
Trip+: Benchmarking Agents in Personalized Interactive Travel Planning
Junle Chen, Wei Chen, Yehong Xu +6
Interactive travel planning has become a popular use case for language models. Agents are deployed to manage evolving preferences and unexpected disruptions over multiple turns. Su…
MExam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions
Zhengjun Huang, Wenxuan Liu, Zhoujin Tian +6
Language agents are increasingly deployed over accumulating multimodal information, yet existing benchmarks assume a human-human form with sparse visuals and straightforward conten…
LifeSide: Benchmarking Agents as Lifelong Digital Companions
Yuqian Wu, Zhijie Deng, Wei Chen +8
Lifelong digital companions must integrate cross-session cues, continually update their understanding of users, and adapt to shifting privacy boundaries. Existing evaluations fail…
LiCoMemory: Lightweight and Cognitive Agentic Memory for Efficient Long-Term Reasoning
Zhengjun Huang, Zhoujin Tian, Qintian Guo +5
Large Language Model (LLM) agents exhibit remarkable conversational and reasoning capabilities but remain constrained by limited context windows and the lack of persistent memory.…
EraRAG: Efficient and Incremental Retrieval Augmented Generation for Growing Corpora
Fangyuan Zhang, Zhengjun Huang, Yingli Zhou +6
Graph-based Retrieval-Augmented Generation (Graph-RAG) enhances large language models (LLMs) by structuring retrieval over an external corpus. However, existing approaches typicall…