activity
20242026
collaborators

5 papers

cs.AI2026

Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning

Zhengbo Zhang, Changtao Miao, Jinbo Su +10

Multimodal large language models (MLLMs) have demonstrated impressive capabilities in many visual tasks, but they often struggle with factual grounding when confronted with complex…

cs.CV2025

AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models

Shixiong Xu, Chenghao Zhang, Lubin Fan +5

Large visual language models (LVLMs) have demonstrated impressive performance in coarse-grained geo-localization at the country or city level, but they struggle with fine-grained s…

cs.CV2025

Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models Stronger

Qi Yang, Chenghao Zhang, Lubin Fan +3

Recent advancements in Large Vision Language Models (LVLMs) have significantly improved performance in Visual Question Answering (VQA) tasks through multimodal Retrieval-Augmented…

cs.CV2024

A Survey of Low-shot Vision-Language Model Adaptation via Representer Theorem

Kun Ding, Ying Wang, Gaofeng Meng +1

The advent of pre-trained vision-language foundation models has revolutionized the field of zero/few-shot (i.e., low-shot) image recognition. The key challenge to address under the…

cs.CV2024

Calibrated Cache Model for Few-Shot Vision-Language Model Adaptation

Kun Ding, Qiang Yu, Haojian Zhang +2

Cache-based approaches stand out as both effective and efficient for adapting vision-language models (VLMs). Nonetheless, the existing cache model overlooks three crucial aspects.…