Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
MExam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions
Zhengjun Huang, Wenxuan Liu, Zhoujin Tian +6
Language agents are increasingly deployed over accumulating multimodal information, yet existing benchmarks assume a human-human form with sparse visuals and straightforward conten…
cs.CL2026
LifeSide: Benchmarking Agents as Lifelong Digital Companions
Yuqian Wu, Zhijie Deng, Wei Chen +8
Lifelong digital companions must integrate cross-session cues, continually update their understanding of users, and adapt to shifting privacy boundaries. Existing evaluations fail…