2 papers
cs.CV2025
A Simple-but-effective Baseline for Training-free Class-Agnostic Counting
Yuhao Lin, Haiming Xu, Lingqiao Liu +1
Class-Agnostic Counting (CAC) seeks to accurately count objects in a given image with only a few reference examples. While previous methods achieving this relied on additional trai…
cs.CV2024
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning
Hai-Ming Xu, Qi Chen, Lei Wang +1
Recent advancements in Multimodal Large Language Models (MLLMs) have generated significant interest in their ability to autonomously interact with and interpret Graphical User Inte…