Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees
Yuankai Li, Tinghui Zhu, Ha Min Son +3
Large Multimodal Models (LMMs) have recently emerged as promising backbones for GUI-agent models, where high-resolution GUI screenshots are introduced to the prompts at each iterat…
cs.AI2024
Dissecting Dissonance: Benchmarking Large Multimodal Models Against Self-Contradictory Instructions
Jin Gao, Lei Gan, Yuankai Li +2
Large multimodal models (LMMs) excel in adhering to human instructions. However, self-contradictory instructions may arise due to the increasing trend of multimodal interaction and…