collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

Asymmetric Cross-Modal Fine-Grained Visual Categorization: ACF-Net and the BirdPro Benchmark

Bohan Deng, Shuo Ye, Zitong Yu

Audio-visual cross-modal Fine-Grained Visual Categorization (FGVC) aims to identify fine-grained categories by jointly leveraging visual and auditory information. However, FGVC und…

cs.CV2026

DeceptionX: From Multimodal Evidence to Explainable Deception Detection

Jiayu Zhang, Shuo Ye, Jiajian Huang +8

Deception detection is a critical and highly challenging task within affective computing and behavioral analysis. Existing deep learning methods typically treat this task as a stra…

cs.CV2026

PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models

Zihan Song, Shuo Ye, Bo Zhao +4

Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindranc…

cs.CV2026

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning

Yelin Wang, Zijia Song, Shuo Ye +6

Remote Sensing Image Change Captioning (RSICC) aims to describe changes between bi-temporal remote sensing images and holds significant research and application value. However, mos…

cs.CV2026

Text-Guided Multimodal Unified Industrial Anomaly Detection

Zewen Li, Shuo Ye, Zitong Yu +2

Industrial anomaly detection based on RGB-3D multimodal data has emerged as a mainstream paradigm for intelligent quality inspection. However, existing unsupervised methods suffer…

cs.CV2026

Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent Purification

Jiayu Zhang, Shuo Ye, Qilang Ye +3

Recent Audio-Visual Question Answering (AVQA) methods have advanced significantly. However, most AVQA methods lack effective mechanisms for handling missing modalities, suffering f…