3 papers
cs.CL2026
Learning to Retrieve: Dual-Level Long-Term Memory for Text-to-SQL Agents
Yibo Wang, Nikki Lijing Kuang, Philip S. Yu +2
Interactive text-to-SQL agents solve database tasks through multi-turn interactions involving schema exploration, query execution, feedback interpretation, and decision revision. L…
cs.CL2026
MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization
Weizhi Zhang, Xiaokai Wei, Wei-Chieh Huang +4
Recent advancements in Large Language Models (LLMs) have expanded context windows to million-token scales, yet benchmarks for evaluating memory remain limited to short-session synt…
cs.CL2025
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
Li Li, Peilin Cai, Ryan A. Rossi +21
We present PersonaConvBench, a large-scale benchmark for evaluating personalized reasoning and generation in multi-turn conversations with large language models (LLMs). Unlike exis…