9 papers
REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot Alignment
Kai Ye, Xianwei Mao, Sheng Zhou +6
Knowledge-intensive Visual Question Answering (KI-VQA) frequently suffers from severe knowledge conflicts caused by the inherent limitations of open-domain retrieval. However, exis…
Towards Scalable Lightweight GUI Agents via Multi-role Orchestration
Ziwei Wang, Junjie Zheng, Leyang Yang +7
Autonomous Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) enable digital automation on end-user devices. While scaling both parameters an…
Generalizable Self-Evolving Memory for Automatic Prompt Optimization
Guanbao Liang, Yuanchen Bei, Sheng Zhou +5
Automatic prompt optimization is a promising approach for adapting large language models (LLMs) to downstream tasks, yet existing methods typically search for a specific prompt spe…
MaS-VQA: A Mask-and-Select Framework for Knowledge-Based Visual Question Answering
Xianwei Mao, Kai Ye, Sheng Zhou +4
Knowledge-based Visual Question Answering (KB-VQA) requires models to answer questions by integrating visual information with external knowledge. However, retrieved knowledge is of…
Towards Scalable Web Accessibility Audit with MLLMs as Copilots
Ming Gu, Ziwei Wang, Sicen Lai +3
Ensuring web accessibility is crucial for advancing social welfare, justice, and equality in digital spaces, yet the vast majority of website user interfaces remain non-compliant,…
Leveraging LLM Agents for Automated Video Game Testing
Chengjia Wang, Lanling Tang, Ming Yuan +3
Testing MMORPGs (Massively Multiplayer Online Role-Playing Games) is a critical yet labor-intensive task in game development due to their complexity and frequent updating nature. T…