33 citations · 36 across the 3 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
Shijie Zhou, Viet Dac Lai, Hao Tan +4
Graphical user interface (GUI) grounding is a key capability for computer-use agents, mapping natural-language instructions to actionable regions on the screen. Existing Multimodal…
cs.CV2025
When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification
Zirui Pang, Haosheng Tan, Yuhan Pu +4
Image classification benchmark datasets such as CIFAR, MNIST, and ImageNet serve as critical tools for model evaluation. However, despite the cleaning efforts, these datasets still…
cs.CV2021★ 33 cited
VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning
Hao Tan, Jie Lei, Thomas Wolf +1
Video understanding relies on perceiving the global content and modeling its internal connections (e.g., causality, movement, and spatio-temporal correspondence). To learn these in…