From the 2 of 32 linked papers with an AI index.
32 papers
Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
Zheng Wu, Chenhao Xue, Shijie Zheng +3
The paper identifies a "salience bias" in large language models where explicit but irrelevant details cause the models to overlook implicit commonsense knowledge, and shows that th…
Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation
Zheng Wu, Yibo Luo, Pu Zhang +2
The paper introduces ESPP, a three-stage evaluation framework that uses a panel of diverse, evidence‑grounded personas to rate generative UI screenshots, improving alignment with h…
Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents
Pengzhou Cheng, Haowen Hu, Zheng Wu +4
Graphical user interface (GUI) agents powered by multimodal large language models (MLLMs) have shown greater promise for human-interaction. However, due to the high fine-tuning cos…
MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft
Tianjie Ju, Yueqing Sun, Zheng Wu +7
Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to sustain exploration in dynamic op…
EVA: Evolving Semantic Adversaries for Red-Teaming GUI Agents Against Environmental Injection Attacks
Yijie Lu, Manman Zhao, Tianjie Ju +6
Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) are increasingly deployed yet vulnerable to Environmental Injection Attacks (EIAs).However…
Mobile-Aptus: Confidence-Driven Proactive and Robust Interaction in MLLM-based Mobile-Using Agents
Zheng Wu, Pengzhou Cheng, Zongru Wu +5
Recent advancements in multimodal large language models (MLLMs) have shown exceptional potential in enabling mobile-using agents to autonomously execute human instructions. However…