2 papers
cs.CV2026
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
Weiliang Chen, Haowen Sun, Jun Gao +40
A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, w…
cs.CL2026
QUMem: Personalized Memory for Query-Conditioned User-State Inference in LLM Agents
Heng Wang, Yifei Li, Lingling Zhang +4
Large language model (LLM) agents increasingly use external memory systems to support personalization by drawing on long and evolving interaction histories, in which user preferenc…