2 papers
cs.AI2026
PERMA: Benchmarking Personalized Memory Agents via Event-Driven Preference and Realistic Task Environments
Shuochen Liu, Junyi Zhu, Long Shu +11
Empowering large language models with long-term memory is crucial for building agents that adapt to users' evolving needs. Existing evaluations of this capability typically interle…
cs.CL2025
Chart-HQA: A Benchmark for Hypothetical Question Answering in Charts
Xiangnan Chen, Yuancheng Fang, Qian Xiao +5
Multimodal Large Language Models (MLLMs) have garnered significant attention for their strong visual-semantic understanding. Most existing chart benchmarks evaluate MLLMs' ability…