6 papers
Multi-modal Test-time Adaptation via Adaptive Probabilistic Gaussian Calibration
Jinglin Xu, Yi Li, Chuxiong Sun +3
Multi-modal test-time adaptation (TTA) enhances the resilience of benchmark multi-modal models against distribution shifts by leveraging the unlabeled target data during inference.…
CausalFSFG: Rethinking Few-Shot Fine-Grained Visual Categorization from Causal Perspective
Zhiwen Yang, Jinglin Xu, Yuxin Pen
Few-shot fine-grained visual categorization (FS-FGVC) focuses on identifying various subcategories within a common superclass given just one or few support examples. Most existing…
DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding
Geng Li, Jinglin Xu, Yunzhen Zhao +1
Humans can effortlessly locate desired objects in cluttered environments, relying on a cognitive mechanism known as visual search to efficiently filter out irrelevant information a…
Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
Hulingxiao He, Geng Li, Zijun Geng +2
Multi-modal large language models (MLLMs) have shown remarkable abilities in various visual understanding tasks. However, MLLMs still struggle with fine-grained visual recognition…
CountMamba: Exploring Multi-directional Selective State-Space Models for Plant Counting
Hulingxiao He, Yaqi Zhang, Jinglin Xu +1
Plant counting is essential in every stage of agriculture, including seed breeding, germination, cultivation, fertilization, pollination yield estimation, and harvesting. Inspired…
SIA-OVD: Shape-Invariant Adapter for Bridging the Image-Region Gap in Open-Vocabulary Detection
Zishuo Wang, Wenhao Zhou, Jinglin Xu +1
Open-vocabulary detection (OVD) aims to detect novel objects without instance-level annotations to achieve open-world object detection at a lower cost. Existing OVD methods mainly…