Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
StreamMemBench: Streaming Evaluation of Agent Memory for Future-Oriented Assistance
Guanming Liu, Yuqi Ren, Hansu Gu +5
A central role of personal-agent memory is to turn stored information and prior interactions into future-oriented assistance. In daily use, useful cues come from what the agent obs…
cs.AI2026
MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
Yanxu Zhu, Shitong Duan, Xiangxu Zhang +7
Recently Multimodal Large Language Models (MLLMs) have achieved considerable advancements in vision-language tasks, yet produce potentially harmful or untrustworthy content. Despit…
cs.AI2025
Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values
Jing Yao, Xiaoyuan Yi, Shitong Duan +8
As Large Language Models (LLMs) achieve remarkable breakthroughs, aligning their values with humans has become imperative for their responsible development and customized applicati…