activity
20242026
collaborators

9 papers

cs.CV2026

RGBX-R1: Visual Modality Chain-of-Thought Guided Reinforcement Learning for Multimodal Grounding

Jiahe Wu, Bing Cao, Qilong Wang +3

Multimodal Large Language Models (MLLM) are primarily pre-trained on the RGB modality, thereby limiting their performance on other modalities, such as infrared, depth, and event da…

cs.CV2025

AM-Net: Adaptively Aligned Multi-Scale Moment for Few-Shot Action Recognition

Zilin Gao, Qilong Wang, Bingbing Zhang +2

Thanks to capability to alleviate the cost of large-scale annotation, few-shot action recognition (FSAR) has attracted increased attention of researchers in recent years. Existing…

cs.CV2025

Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models

Xiaojie Yin, Qilong Wang, Qinghua Hu

Vision-language models (VLMs) pre-trained on web-scale data exhibit promising zero-shot generalization but often suffer from semantic misalignment due to domain gaps between pre-tr…

cs.CV2025

TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action Recognition

Yilong Wang, Zilin Gao, Qilong Wang +3

Going beyond few-shot action recognition (FSAR), cross-domain FSAR (CDFSAR) has attracted recent research interests by solving the domain gap lying in source-to-target transfer lea…

cs.CV2025

DALIP: Distribution Alignment-based Language-Image Pre-Training for Domain-Specific Data

Junjie Wu, Jiangtao Xie, Zhaolin Zhang +4

Recently, Contrastive Language-Image Pre-training (CLIP) has shown promising performance in domain-specific data (e.g., biology), and has attracted increasing research attention. E…

cs.CV2025

CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity Quantification

Wenlong Yu, Qilong Wang, Chuang Liu +2

Explainability is a critical factor influencing the wide deployment of deep vision models (DVMs). Concept-based post-hoc explanation methods can provide both global and local insig…