activity
20242026
collaborators

7 papers

cs.CV2026

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception

Geng Li, Yuxin Peng

While Multimodal Large Language Models (MLLMs) demonstrate impressive general capabilities, they struggle with fine-grained perception in ultra-high-resolution (UHR) images, partic…

cs.CV2026

LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal Observations

Huangbiao Xu, Huanqi Wu, Xiao Ke +1

Real-world multimodal learning is often hindered by missing modalities. While Incomplete Multimodal Learning (IML) has gained traction, existing methods typically rely on the unrea…

cs.CV2025

CausalFSFG: Rethinking Few-Shot Fine-Grained Visual Categorization from Causal Perspective

Zhiwen Yang, Jinglin Xu, Yuxin Pen

Few-shot fine-grained visual categorization (FS-FGVC) focuses on identifying various subcategories within a common superclass given just one or few support examples. Most existing…

cs.CV2025

DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding

Geng Li, Jinglin Xu, Yunzhen Zhao +1

Humans can effortlessly locate desired objects in cluttered environments, relying on a cognitive mechanism known as visual search to efficiently filter out irrelevant information a…

cs.CV2025

Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models

Hulingxiao He, Geng Li, Zijun Geng +2

Multi-modal large language models (MLLMs) have shown remarkable abilities in various visual understanding tasks. However, MLLMs still struggle with fine-grained visual recognition…

cs.CV2024

CountMamba: Exploring Multi-directional Selective State-Space Models for Plant Counting

Hulingxiao He, Yaqi Zhang, Jinglin Xu +1

Plant counting is essential in every stage of agriculture, including seed breeding, germination, cultivation, fertilization, pollination yield estimation, and harvesting. Inspired…