3 papers
cs.CL2026
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
Xintao Wang, Jian Yang, Weiyuan Li +8
Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning and generation, serving as the foundation for advanced persona simulation and Role-Playing Langu…
cs.AI2026
LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
Zihao Cheng, Weixin Wang, Yu Zhao +15
Long-term memory is fundamental for personalized agents capable of accumulating knowledge, reasoning over user experiences, and adapting across time. However, existing memory bench…
cs.CL2025
SI-Bench: Benchmarking Social Intelligence of Large Language Models in Human-to-Human Conversations
Shuai Huang, Wenxuan Zhao, Jun Gao
As large language models (LLMs) develop anthropomorphic abilities, they are increasingly being deployed as autonomous agents to interact with humans. However, evaluating their perf…