2 papers
cs.CV2026
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
Shijie Zhou, Viet Dac Lai, Hao Tan +4
Graphical user interface (GUI) grounding is a key capability for computer-use agents, mapping natural-language instructions to actionable regions on the screen. Existing Multimodal…
cs.CV2025
When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification
Zirui Pang, Haosheng Tan, Yuhan Pu +4
Image classification benchmark datasets such as CIFAR, MNIST, and ImageNet serve as critical tools for model evaluation. However, despite the cleaning efforts, these datasets still…