benchmark 1benchmark dataset 1commonsense reasoning 1generative ui evaluation 1human-computer interaction 1large language models 1llm judging 1longitudinal analysis 1medical visual question answering 1multimodal language models 1persona panels 1prompt engineering 1
From the 3 of 3 linked papers with an AI index.
3 papers
cs.CL2026
Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
Zheng Wu, Chenhao Xue, Shijie Zheng +3
The paper identifies a "salience bias" in large language models where explicit but irrelevant details cause the models to overlook implicit commonsense knowledge, and shows that th…
cs.CL2026
Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation
Zheng Wu, Yibo Luo, Pu Zhang +2
The paper introduces ESPP, a three-stage evaluation framework that uses a panel of diverse, evidence‑grounded personas to rate generative UI screenshots, improving alignment with h…
cs.CV2026
LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA
Zhilin Wu, Zhangkai Ni, Chengmei Yang +4
The paper introduces LoMeVQA, a large benchmark of 206K longitudinal medical visual question answering pairs designed to evaluate temporal reasoning over sequential medical images,…