21 papers
SocialPersona: Benchmarking Personalized Profiling and Response with Multimodal Social-Media Context
Qinkai Zhang, Yanyan Zhao, Xin Lu +3
Personalized language-model assistants are often evaluated through a memory lens: can a model recall preferences users have explicitly stated in dialogue? More comprehensive person…
Large Language Model Agents Are Not Always Faithful Self-Evolvers
Weixiang Zhao, Yingshuo Wang, Yichen Zhang +5
Self-evolving large language model (LLM) agents continually improve by accumulating and reusing past experience, yet it remains unclear whether they faithfully rely on that experie…
ENPMR-Bench: Benchmarking Proactive Memory Retrieval for Emotional Support Agents
Xing Fu, Yulin Hu, Mengtong Ji +5
Memory-augmented language agents are increasingly deployed in affective applications such as emotional support, where understanding and responding to users' latent emotional needs…
When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents
Jiahe Guo, Xiangran Guo, Yulin Hu +8
Long-term memory enables large language model (LLM) agents to support personalized and sustained interactions. However, most work on personalized agents prioritizes utility and use…
TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
Xingyu Sui, Yanyan Zhao, Yulin Hu +3
Emotional Support Conversation requires not only affective expression but also grounded instrumental support to provide trustworthy guidance. However, existing ESC systems and benc…
ConflictBench: Evaluating Human-AI Conflict via Interactive and Visually Grounded Environments
Weixiang Zhao, Haozhen Li, Yanyan Zhao +5
As large language models (LLMs) evolve into autonomous agents capable of acting in open-ended environments, ensuring behavioral alignment with human values becomes a critical safet…